Papers for

automated research teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Autonomous AI improves itself safely with structured experiments

RSI-Master: Structuring Experiments to Guide Autonomous Model Improvement

Abstract: Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit open-ended experimental actions through hacking, and repeated experimentation may lead to strategy lock-in, where an early direction is refined rather than reconsidered. We introduce RSI-Master, which addresses the two challenges at two levels: regularize step-wise actions, avoiding hacking behaviors, and promote well-structured exploration of research directions, avoiding strategy lock-in. RSI-Master consists of an Experiment OS, which enables regularized experimental actions and maintains persistent, traceable experimental records, and Reviewer-Guided Research Orchestration, which organizes Workers and Reviewers in a dynamically growing research DAG. Workers explore diverse research directions and Reviewers compare evidence across related experiments for subsequent explorations. On PostTrainBench with Qwen3-4B-Base, it averages 54.49 versus 46.53 for the strongest agent baseline, with a 0.0\% hacking rate. Scaling to 35B model, RSI-Master surpasses the human-developed Instruct model on LiveCodeBench-v6 (41.21 vs. 37.36) and SciCode, and reaches a nonzero score on HorizonMath, a benchmark of unsolved research problems on which most frontier models score near zero.

Mon 28 SeptArtificial Intelligence
The gist
Improving AI models by letting them experiment on their own can cause problems like cheating or sticking to early ideas without trying new ones. The authors introduce RSI-Master, a system that organizes and regulates AI experiments to avoid these problems. It keeps detailed records and uses a team-like approach where different AI parts test ideas and review results to find better directions. This method outperforms other automated approaches and even beats some human-designed models on difficult tests.
Open → 2609.35561v1

Lara language helps machines check scientific claims in seconds

Beyond Natural Language: An Agent-Native Language for Autonomous Science

Abstract: As autonomous AI agents take on every stage of scientific inquiry, research output is expanding far beyond human review capacity. Yet scientific communication still relies on natural-language prose: an informal medium prone to ambiguity, hidden assumptions, and untracked limitations that machines cannot reliably audit. We introduce Lara, a machine-checkable language and protocol for checking and revising support for research claims. By turning research arguments into executable artifacts, Lara provides an epistemic kernel for autonomous science: it enables automated validation pipelines for research agents, lets declared bridges connect arguments across papers into an auditable network, and allows both humans and machines to recheck the standing of an encoded claim in milliseconds. In a Lara program, authors explicitly declare their claims, supporting evidence and assumptions, and known objections or limitations. A lightweight, deterministic checker adjudicates these interactions, assigning each claim a reproducible status: "justified", "defeated", "contested", or "gap", which marks a claim whose support is incomplete and locates the unanswered question. Case studies cover empirical review, a philosophical debate without measurements, and the loss of support when an assumed axiom is withdrawn. We establish the metatheory of claim checking and cross-context argument transport, and mechanize the semantic guarantees in Lean 4 (roughly 117,000 lines), leaving three arguments on paper. The audited public metatheory is "sorry"-free and uses only Lean's three standard axioms; some executable examples additionally trust native evaluation.

Mon 21 SeptProgramming LanguagesArtificial Intelligence
The gist
Scientific research is growing so fast that humans cannot review all new findings. The authors created Lara, a special language that turns scientific claims and their support into computer-checkable programs. This lets machines quickly validate if claims are justified, challenged, or need more proof. Lara also connects claims across different papers, making it easier to trace ideas and spot gaps. This tool could help machines and humans keep scientific knowledge clear and trustworthy.
Open → 2609.25421v1