Papers for

automated research tool builders

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Scientific agents show weak robustness to errors in multi-step problem solving

The Marathon of Scientific Reasoning: Robustness of Scientific Agents to Perturbations in Multi-Turn Interactions

Abstract: Large language model (LLM)-based scientific agents are increasingly used for scientific problem solving, yet their robustness to imperfections arising during multi-turn interactions remains poorly understood. We introduce \textsc{SciARP} (\textbf{Sci}entific \textbf{A}gent \textbf{R}obustness to \textbf{P}erturbations), a benchmark for evaluating scientific agents under scientifically plausible perturbations throughout multi-turn problem solving. \textsc{SciARP} transforms 620 scientific problems into interdependent tasks of 3--13 turns and defines 13 perturbation types spanning problem understanding, evidence processing, reasoning, and conclusion formation. Clean and perturbed versions of each task are independently executed under matched settings, producing paired live trajectories for evaluating both task success and process reliability. Experiments across eight LLMs from four model families reveal three key robustness characteristics. First, different classes of scientific perturbations exhibit distinct robustness profiles and can decouple task progression from scientific reliability: agents may continue advancing through the task even after their information or reasoning has become unreliable. Second, stronger clean-task performance does not necessarily translate into stronger robustness, as models with higher clean-task accuracy can exhibit larger degradation under perturbation. Third, perturbation effects exhibit strong temporal dynamics: they may remain latent for multiple turns before emerging and subsequently propagate through downstream dependencies. Together, these findings show that current scientific agents remain insufficiently robust to scientifically plausible perturbations, with failures often remaining undetected, propagating, and resisting recovery.

Mon 28 SeptArtificial Intelligence
The gist
Large language models that help solve scientific problems can make mistakes, especially when dealing with several steps of reasoning. The authors created a test to see how these agents handle different scientific challenges and errors over multiple steps. They found that mistakes can go unnoticed as the agent keeps working, and stronger performance on easy tasks doesn't guarantee staying reliable under difficult conditions. Problems with the agent's thinking can build up over time and spread through later steps. This means current systems still struggle to stay accurate and reliable when real-world issues appear during complex scientific reasoning.
Open → 2609.34537v1