Papers for

enterprise data engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Semantic indexing system improves context-aware knowledge extraction

5W1H+Which: Context-Valid Semantic Indexing with Progressive Ontology Binding

Abstract: Transforming raw data into queryable knowledge requires both early extraction of reusable information and explicit types, relations, and applicability conditions for particular tasks. If indexing selects content too early around a single business schema, later tasks may be unable to use information that was omitted. If the index retains only open-ended text, however, rule-based reasoning lacks checkable premises. We propose 5W1H+Which, a semantic indexing design that separates content extraction from ontology binding. The 5W1H questions organize source-grounded content units; Which points to versioned ontology elements and records mapping relations, scope, and validation status. Time, location, system environment, and participant roles are not merely retrieval labels: together, they constrain the contexts in which facts, bindings, and rules apply. Unbound content remains searchable, while bound content enters a formal reasoning path only after premise checks. The method further distinguishes business valid time, system knowledge time, and operational traces, and uses dependency records to support binding revalidation and the maintenance of derived conclusions. A worked example of migration from an on-premises server to a cloud environment illustrates the different treatment of world-state changes, ontology-version changes, and changes in rule applicability. We formulate three groups of falsifiable hypotheses concerning cross-task evidence coverage, control of contextual misuse, and incremental update cost. The planned evaluation includes a strong typed fact-graph baseline with the same evidence, temporal information, and budget, to test whether benefits arise from 5W1H organization, deferred binding, or additional information and engineering effort. The contribution is a testable indexing mechanism, not a claim to a new universal ontology or a demonstrated performance advantage.

Mon 28 SeptArtificial IntelligenceComputation and LanguageInformation Retrieval
The gist
Turning raw data into useful knowledge requires both finding important facts early and knowing exactly what those facts mean. The authors propose a method called 5W1H+Which that first gathers raw facts by answering who, what, when, where, why, and how questions, and then links those facts to clear definitions in an organized way. This method keeps track of different times and settings to ensure facts are used correctly in context, while also allowing flexible searching. They test it using an example of moving servers to the cloud and set up ways to measure if their method helps later tasks.
Open → 2609.35184v1

Enterprise data agents improve reasoning with learned identity routing

Learned Enterprise Data Comprehension: Compression and Routing for Data Agents

Abstract: Structured-data agents in enterprise settings must reason over complex data environments whose relevant evidence is distributed across schemas, relationships, policies, and recurring business roles. Modern agentic systems often address this burden through reusable markdown-style memory or skill files that preserve previously discovered information for later queries, reducing the need to rediscover the same structure repeatedly. This is useful, but it obscures a natural division of labor: agents are well suited to semantic reasoning, while learned systems are well suited to predicting and organizing recurring structure. We introduce latent equivalence learning to bridge this gap. The framework separates persistent task-relevant identities from their dataset-relative realizations. In our realization, supporting and opposing evidence shape support-realized Gaussian prototypes that learn how those identities are expressed in a particular data environment, while soft-membership profiles retain distinctions lost under a hard assignment. A separate learned query-prototype system represents recurring evidential requirements and maps them through a learned compatibility function into the same persistent identity structure. This identity-factorized, query-conditioned routing materializes the relevant dataset-specific evidence for downstream reasoning, allowing the agent to operate over an already organized evidential state rather than reconstructing cross-schema structure at every query. On the Data Agent Benchmark, spanning 54 queries across 12 heterogeneous datasets, our full implementation achieves 94.67% dataset-macro stratified Pass@1 over five complete trials and 258/270 successful raw query attempts, compared with 55.51% for the benchmark's Claude Opus 4.6 reference agent, ranking first among 40 leaderboard entries at submission.

Mon 21 SeptArtificial Intelligence
The gist
Enterprise data agents have trouble finding and using related pieces of information spread across complex databases. The authors created a method that helps the agent understand and organize important identities in different datasets, so it can quickly find relevant evidence without starting from scratch every time. This method uses learned patterns to link identities and queries, improving the agent's ability to reason with structured data. Their approach scored much higher than previous systems on a challenging data benchmark.
Open → 2609.25286v1

Graph databases differ widely in speed and costs for queries and updates

Graph Memory for LLM Agents: At What Cost? A Comparative Evaluation of Query, Ingest, and Update Performance Across Graph Database Engines

Abstract: Graph databases are frequently positioned as categorically necessary for connected-data workloads, yet the systems dimension along which they actually differ - query planning, indexing, and data-readiness cost - is rarely isolated from vendor framing. We construct a synthetic, biomedical-shaped property graph (1.02 million nodes, 5.34 million total node and edge rows) and a twenty-query workload spanning neighborhood lookups, bounded paths, set intersections, anti-joins, grouped aggregation, top-k ranking, temporal filters, full scans, and relational joins. We benchmark Corvic AI - a purpose-built columnar query engine underlying Corvic's ontology management layer ("memories")- against seven purpose-built or graph-extension database systems (LoraDB, Ladybug, DuckPGQ, Memgraph, Neo4j, HugeGraph, and FalkorDB) at three graph scales spanning three orders of magnitude. We report query latency geomeans, bulk-ingest throughput, point-update latency, and answer correctness for each system, and we derive a simple total-cost-of-ownership model that expresses the ingest/query trade-off as a function of query volume. Our central finding is that no system in this sample is categorically fastest: a native graph engine (Ladybug) outperforms Corvic AI on narrow, bounded-neighborhood shapes, while Corvic AI is faster on shapes that scan or join a large fraction of the graph, and a system implementing graph query syntax via SQL/PGQ (DuckPGQ) is measurably slower purely due to query-plan choice. The dominant cost differential in our data is not query latency but the cost of making data queryable at all: bulk-ingest throughput varies by three orders of magnitude across engines (5.0k-4.3M rows/s), a gap that a simple crossover-point calculation shows dominates total cost for any workload with fewer than roughly 105 queries per data refresh.

Sun 20 SeptDatabasesArtificial Intelligence
The gist
Not all graph databases work equally well for searching and updating connected data, even if they seem similar on the surface. The authors tested eight different database systems using a large biomedical network and a variety of common query types. They found that some systems were faster for small, local searches, while others were better for big scans or complex joins. Most importantly, the time it takes to load the data into the database varied hugely and often dominated the overall cost for tasks with fewer queries. This means choosing the right system depends a lot on how often you query versus how often you update your data.
Open → 2609.23315v1

Self demonstrations boost database schema matching with LLMs

Surprising Effectiveness of Self-Demonstrations in Enhancing Schema-Ontology Mapping with LLMs

Abstract: Integrating heterogeneous relational databases into a centralized ontology remains a persistent challenge in enterprise knowledge representation, primarily due to semantic heterogeneity, cryptic schema naming, missing metadata, and the abstraction gap between relational schemas and ontological models. Although large language models (LLMs) offer strong semantic reasoning capabilities, we show that directly applying them through one-shot prompting or naive multi-stage pipelines leads to poor performance for schema-ontology mapping. This paper presents a self-demonstration-driven approach that combines a neuro-symbolic task decomposition with a novel mechanism for automatically generating pattern-guided, dependency-aware demonstrations to address this integration challenge. Our approach incorporates two key strategies to achieve substantial accuracy gains over existing LLM-based schema integration methods: (i) a neuro-symbolic decomposition of the task into cascaded sub-tasks, where symbolic constraints structure the search space and LLMs perform semantic reasoning within each focused sub-task, and (ii) self-generated demonstrations guided by domain-agnostic patterns to supervise each sub-task. Experiments on three of the most challenging scenarios from the RODI benchmark show that our approach achieves state-of-the-art performance, substantially outperforming (25 percentage points F1 improvements) both traditional schema-to-ontology mapping techniques and recent LLM-based schema-to-ontology and schema matching approaches. Ablation studies further reveal the significant benefits of pattern-guided self-demonstrations and the complementary benefits of neuro-symbolic task decomposition.

Sat 12 SeptArtificial Intelligence
The gist
Combining data from different databases into one knowledge model is hard because database layouts can be confusing and incomplete. The authors found that simply asking large language models (LLMs) to do this mapping often fails. They created a method that breaks the task into smaller steps with clear rules and automatically generates helpful examples to guide the LLMs. This method works much better than previous attempts, improving accuracy by a large margin on standard tests.
Open → 2609.13776v1

Federated knowledge graphs enable multi-hop question answering without sharing raw data

FedV-KGQA in Practice: Design Lessons and an Interactive Prototype

Abstract: Knowledge graph question answering usually assumes that one system can reach the whole graph. In practice, facts are often held by organizations that share entity identifiers but own disjoint relation types, so no single party sees a complete reasoning chain. This poster presents the empirical findings of FedV-KGQA on multi-hop question answering over such vertically partitioned graphs. Each silo enriches its local graph and trains a knowledge graph embedding on its own triples. A server then concatenates the silo-specific entity views, anchors the projected question at the topic entity, and ranks candidates by similarity. Raw triples and relation embeddings never leave a silo. Comparing the FedV-KGQA experiments with one another yields three results. First, federated fusion recovers most of the centralized accuracy, while a single silo recovers little. Second, anchoring and enrichment matter more than the choice of embedding model. Third, the cheapest encoder depends on the target accuracy rather than on parameter count. This poster paper contributes that cross-experiment comparison, four design lessons drawn from it, and an interactive prototype that runs real inference and traces the full pipeline, per question, on released checkpoints.

Sat 12 SeptArtificial IntelligenceComputation and LanguageInformation Retrieval
The gist
Question answering systems that use knowledge graphs usually need access to all the information in one place. The authors study what happens when different organizations hold parts of the knowledge graph with unique relationships but share common entity identifiers. They show that combining knowledge locally enriched and embedded in separate silos can nearly match the accuracy of a system that sees everything. They also find that focusing on linking questions to key entities and enriching the data matters more than using complex embedding models. The authors provide lessons from comparing different settings and an interactive prototype that demonstrates the approach.
Open → 2609.13661v1