Deep research agents use rubric feedback to improve process and results

Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents

Computation and LanguageArtificial Intelligence

Summary

Many AI systems learn by getting rewarded only for their final answers, which ignores how well their intermediate steps are done. The authors propose a new method that uses detailed rubrics to give feedback not just on the final answer but on each step along the way. This helps the AI better understand which parts of its process helped meet the task requirements. Their approach, Dr.Credit, improves how AI research agents gather evidence and produce reports, performing better than previous systems on multiple benchmarks. The method works well even with limited resources, showing promise for a range of tasks judged by rubrics.

What this means in practice

  • For ai research teams: Improve training of AI agents that perform complex research tasks by providing feedback on intermediate steps, not just the final output.
  • For natural language processing developers: Develop AI systems that generate higher-quality reports efficiently by using rubric-based process guidance to supervise each research action.

Authors

Yingjian Zhu, Zhenyi Wang, Jiaxin Guo, Kun Ding, Ying Wang, Shen Huang, Xunjie Zhu, Pengjun Xie, Shiming Xiang

Abstract

Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assignment methods rely on ground-truth answers to define process rewards, limiting their applicability to open-ended tasks without canonical solutions. To address this limitation, the proposed rubric-grounded credit uses task requirements as a shared reference for final answer evaluation and process supervision. The information returned by tools is assessed for the additional support it provides toward satisfying each rubric relative to that rubric's history of accepted support. By referencing these histories, credit distinguishes new support from evidence already present in the trajectory while recognizing partial support for each rubric. Dr.Credit uses rubric-grounded credit to supervise intermediate tool turns in an RL framework for deep research agents. The resulting process advantages are combined with GRPO outcome advantages to guide research decisions while retaining supervision of final-report quality. Evaluations on four in-domain and out-of-domain benchmarks show that Dr.Credit outperforms the evaluated open deep research baselines on every primary metric and submetric. Meanwhile, with an 8B-parameter backbone, the trained agent achieves average performance competitive with the evaluated frontier proprietary models. Further analyses suggest more efficient evidence acquisition and higher-quality reports under limited research-turn budgets, motivating the extension of rubric-grounded process supervision to a broader range of rubric-based tasks.