Enterprise data agents improve reasoning with learned identity routing
Learned Enterprise Data Comprehension: Compression and Routing for Data Agents
Artificial Intelligence
Summary
Enterprise data agents have trouble finding and using related pieces of information spread across complex databases. The authors created a method that helps the agent understand and organize important identities in different datasets, so it can quickly find relevant evidence without starting from scratch every time. This method uses learned patterns to link identities and queries, improving the agent's ability to reason with structured data. Their approach scored much higher than previous systems on a challenging data benchmark.
What this means in practice
- •For enterprise data engineers: Enhance query systems to rapidly retrieve relevant data relationships across multiple databases by organizing identities with learned prototypes.
- •For business intelligence teams: Improve automated reasoning over complex business data by using identity-aware routing to reduce repeated structure discovery in queries.
Authors
Ethan Torres, Eric Mills
Abstract
Structured-data agents in enterprise settings must reason over complex data environments whose relevant evidence is distributed across schemas, relationships, policies, and recurring business roles. Modern agentic systems often address this burden through reusable markdown-style memory or skill files that preserve previously discovered information for later queries, reducing the need to rediscover the same structure repeatedly. This is useful, but it obscures a natural division of labor: agents are well suited to semantic reasoning, while learned systems are well suited to predicting and organizing recurring structure. We introduce latent equivalence learning to bridge this gap. The framework separates persistent task-relevant identities from their dataset-relative realizations. In our realization, supporting and opposing evidence shape support-realized Gaussian prototypes that learn how those identities are expressed in a particular data environment, while soft-membership profiles retain distinctions lost under a hard assignment. A separate learned query-prototype system represents recurring evidential requirements and maps them through a learned compatibility function into the same persistent identity structure. This identity-factorized, query-conditioned routing materializes the relevant dataset-specific evidence for downstream reasoning, allowing the agent to operate over an already organized evidential state rather than reconstructing cross-schema structure at every query. On the Data Agent Benchmark, spanning 54 queries across 12 heterogeneous datasets, our full implementation achieves 94.67% dataset-macro stratified Pass@1 over five complete trials and 258/270 successful raw query attempts, compared with 55.51% for the benchmark's Claude Opus 4.6 reference agent, ranking first among 40 leaderboard entries at submission.