Papers for

ai research teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Deep research agents use rubric feedback to improve process and results

Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents

Abstract: Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assignment methods rely on ground-truth answers to define process rewards, limiting their applicability to open-ended tasks without canonical solutions. To address this limitation, the proposed rubric-grounded credit uses task requirements as a shared reference for final answer evaluation and process supervision. The information returned by tools is assessed for the additional support it provides toward satisfying each rubric relative to that rubric's history of accepted support. By referencing these histories, credit distinguishes new support from evidence already present in the trajectory while recognizing partial support for each rubric. Dr.Credit uses rubric-grounded credit to supervise intermediate tool turns in an RL framework for deep research agents. The resulting process advantages are combined with GRPO outcome advantages to guide research decisions while retaining supervision of final-report quality. Evaluations on four in-domain and out-of-domain benchmarks show that Dr.Credit outperforms the evaluated open deep research baselines on every primary metric and submetric. Meanwhile, with an 8B-parameter backbone, the trained agent achieves average performance competitive with the evaluated frontier proprietary models. Further analyses suggest more efficient evidence acquisition and higher-quality reports under limited research-turn budgets, motivating the extension of rubric-grounded process supervision to a broader range of rubric-based tasks.

Mon 28 SeptComputation and LanguageArtificial Intelligence
The gist
Many AI systems learn by getting rewarded only for their final answers, which ignores how well their intermediate steps are done. The authors propose a new method that uses detailed rubrics to give feedback not just on the final answer but on each step along the way. This helps the AI better understand which parts of its process helped meet the task requirements. Their approach, Dr.Credit, improves how AI research agents gather evidence and produce reports, performing better than previous systems on multiple benchmarks. The method works well even with limited resources, showing promise for a range of tasks judged by rubrics.
Open → 2609.34296v1

Language models learn to generate scientific ideas with greater impact

Learning to Ideate for Scientific Impact

Abstract: Scientific ideation is increasingly mediated by large language models, but current ideation systems are usually trained and evaluated on immediately judgeable proxies such as novelty, clarity, and feasibility. This leaves open whether delayed signals of scientific uptake can be used as feedback for steering models toward research directions with higher expected \emph{impact}. We study this question using citation-normalized impact as a noisy but scalable proxy for scholarly uptake. We construct a large-scale dataset from over 100K computer science papers by extracting goal-conditioned idea descriptions and assigning each paper an ordinal, year-normalized citation label. We then train a goal-conditioned reward model to predict citation-impact labels from research goal and idea pairs, and use this reward to align an idea generator through supervised fine-tuning followed by reinforcement learning. To reduce circularity, we evaluate generated ideas with a held-out, reference-grounded protocol that compares model outputs against historical ideas under the same research goal and weights judgments by the reference idea's citation-impact label. Experiments show that our RL-tuned model consistently produces ideas with higher estimated impact than both the base model and supervised fine-tuning baselines. Our findings position scientific impact as a practical, outcome-grounded feedback signal for aligning LLMs in open-ended scientific discovery.

Thu 24 SeptArtificial IntelligenceComputation and Language
The gist
Scientific ideas are hard to judge right away using simple qualities like newness or clarity. The authors studied whether the long-term impact of ideas, as measured by future citations of scientific papers, can guide models to generate better ideas. They trained a system to predict a paper’s future impact from its goals and idea description, then fine-tuned the model to create ideas expected to have higher impact. Their experiments show that this approach helps produce more influential scientific suggestions compared to previous methods.
Open → 2609.29802v1