Enterprise agents struggle with hidden facts in company data questions

Era by Eon: Benchmarking Enterprise Agents on Hidden Knowledge

Software EngineeringArtificial Intelligence

Summary

Some questions about company data have facts that aren't directly stated, making them hard for AI agents to answer. The authors created a test where agents must find answers relying on hidden clues, not just stated facts. They found most AI models handled easy questions well but struggled greatly with these hidden fact questions. This shows current AI tools have trouble understanding subtle or implied information in business settings.

What this means in practice

Authors

Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary

Abstract

In the Era by Eon benchmark, each question states the rules for its answer, and code computes the answer from a generated company's data. When agents can run code, the four strongest models each answer 22 to 25 of 27 such questions, so the benchmark barely separates them. We add eight question templates that depend on hidden facts. No question or document states a hidden fact, and the records that seem to hold it show something else. Other data implies it. For example, the sales system says a customer dropped a purchase because of timing. On a recorded call, the customer blames an outage. For each generated company, code fills each template and computes an exact answer without a language model. We evaluate 12 agents. Each pairs a model with an agent program, which connects it to the company's systems. The best agent answers 18 of its 24 attempts, three per question, correctly. Four of the six models answer at most 6 of 24 with any program. The hardest questions require picking one of several similar records, such as which of three renewal offers a customer signed. All agents together answered two such questions correctly in only 1 of 84 attempts.