Agent platform enforces trustworthy long-running scientific workflows

Research-Native by Construction: Minimal Nodes, Re-verifiable Workflows, and Compounding Memory for Long-Horizon Scientific Agents

Software EngineeringArtificial Intelligence

Summary

It’s hard for AI systems to do long, complex science projects without making mistakes or guessing answers. The authors built a system called AfS that uses strict rules to stop the AI from fudging data or skipping steps. These rules are built into the system’s structure, so mistakes can’t happen without the system catching them. Instead of trusting the AI to behave well, the system makes it impossible to represent bad behavior. This helps ensure research done by the AI is more reliable over many hours and many task runs.

What this means in practice

  • For automation engineers: Build systems that run complex scientific processes unattended for hours with enforced data integrity and traceability.
  • For industrial data teams: Maintain clear, tamper-proof records of multi-step data analyses across extended projects to improve auditability and reduce errors.

Tested on one dataset.

Authors

Di Wang, Yu Liu, Bing Cui, Chaoqun Ji, Dongyuan Ni, Jingyu Lu, Kunlei Cui, Pu Qin

Abstract

We describe AfS (Agent for Science), a platform built for long-horizon scientific work, where a project runs for tens of hours across dozens of agent runs with a human present only occasionally. Most agents for science are general coding agents with a skills folder attached, and they inherit that lineage's failure mode: under pressure to finish, they fabricate, skip, or smooth over. Our design rests on one claim: most of the credibility of machine-made research can be moved from asking the model to behave to making the non-compliant state unrepresentable. We encode research discipline as mechanically enforced laws (commitment before measurement; unforgeable freezing; reports are not facts; evidence persists but verdicts do not; negative results are first-class; mechanical questions to the framework and semantic judgment to the model), organized around three time horizons: a minimal set of research nodes within a run, an inquiry contract with frozen closure conditions and a hash-chained artifact ledger within a project, and a two-tier knowledge base with promotion by rewriting across projects. This is a system description written under one rule: each mechanism appears in exactly one place, with the invariant it enforces, the failure it prevents, the way it is realized, and the cost it imposes. It covers the node contract, the write-path gates, the two-tier memory, and the runtime substrate. Three traces walk real failure attempts through the mechanisms that catch them, and two closed campaigns are included as worked illustrations rather than as an evaluation. We report no benchmark: a process-integrity suite that would support quantitative comparison is under construction, and what we can measure today is only the operating cost of the machinery.