Finance agents get automatic expert-guided evaluation rubrics
FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents
Artificial Intelligence
Summary
Evaluating financial research agents needs careful and expert-based rules to judge their answers accurately. The authors created FinAutoRubric, a system where experts give general guidelines that computer agents use to build detailed, specific rules for checking answers. These rules help write and review financial questions automatically, with human checks for problems. The system was tested on real finance questions and matched expert scores well, even earning preference from analysts in blind tests.
What this means in practice
- •For financial analysts: Generate consistent, customizable evaluation rubrics for assessing financial research tasks across diverse asset classes.
- •For financial technology developers: Build automated review systems that generate and validate assessment criteria dynamically for finance-related AI agents.
Authors
Hoyoung Lee, Suyeol Yun, Jack Haverty, Yunju Cho, Meesong Kim, Daekyung Park, Sumin Kim, Jihoon Kwon, Jasmine Jia Geng, Andrew Chin, Yin Luo, Edward Tong, Yu Yu, Zach Golkhou, Minkyu Kim, Igor Halperin, Young Cha, Alejandro Lopez-Lira, Chanyeol Choi, Yongjae Lee
Abstract
Evaluating finance research agents requires rubrics that reflect expert standards and fix the values correct as of an information cutoff. Expert-reviewed finance benchmarks rely on fixed, per-item rubrics, which are costly to extend and cannot encode each institution's own standard. In FinAutoRubric, experts specify reusable evaluation guidance, while agents and code carry out query-specific rubric generation, review, and validation. This expert guidance governs every agent, as prompts and as rules that code enforces, and a Task Bank of reusable criteria carries it across tasks. In long-horizon loops that follow the expert guidance, a writer agent researches every expected value and a reviewer agent verifies it, and failures escalate to a human. On three expert-authored finance benchmarks, its rubrics track expert scoring as closely as the strongest evaluated generator while stating the expert rubric's expected value for more criteria, their scores agree with human grading, and in-house analysts prefer them in a blind review. The released 100-query FinAutoRubric Benchmark, built from in-house analysts' key questions across 78 tasks and eight asset classes, shows that rubrics from an earlier model generation still leave headroom for a later one.