AI agent swarms make steady progress in scientific research tasks
AI Agent Swarms as Researchers: Progress, Challenges, and Open Questions
Computers and Society
Summary
AI programs working together can now carry out complex research tasks without much human help. The authors gave groups of these AI agents access to scientific papers and tools, with only a simple goal: make real progress. In a few weeks, the agents produced lots of research notes, drafts, and proofs in math and science fields. While not all of it was checked, what was reviewed had no major errors, showing these AI swarms can do routine theoretical research faster than humans. This raises important questions about trusting, reviewing, and crediting work done by AI.
What this means in practice
- •For research software engineers: Use AI agent swarms to accelerate routine theoretical research and generate formal proofs in optimization and physical science domains.
- •For laboratory automation teams: Employ AI-generated experimental proposals to guide and automate preliminary design of new laboratory experiments.
Authors
Sergey Gusev, David E. Bernal Neira
Abstract
Artificial intelligence (AI) agents, language models connected to tools and run in a loop, can now carry out long, multi-step tasks with little supervision. We gave swarms of off-the-shelf coding agents a short statement of scope, from a narrow topic to a whole field, access to the literature and to computing tools, and one standing instruction: make real, correct, useful progress, and do not stop. We supplied no scientific ideas. Within weeks, the agents produced a large body of research notes, paper-length drafts, and formal proofs in five areas of optimization theory and physical science, and proposed untested laboratory experiments in a sixth. We do not claim that all of it is correct or new, but it is not noise: in what we have checked so far, we found no major scientific error, and several results are proved in a proof assistant. The agents produced results faster than we could review them; we estimate that a full review would take us months. Together with two widely discussed 2026 results in mathematics obtained with swarms, our runs suggest that agents can already do a large part of routine theoretical research, at least in areas that we experimented with. This raises questions we cannot yet answer: how to trust results when review, not production, is the scarce resource; what credit and publication counts mean when the human input is a prompt, and why institutions would pay researchers rather than buy computing time; and how people can learn a field, add to what agents do, and stay in control of research they cannot keep up with. Research institutions are not ready: models improve faster than institutions change, so they should decide now how to respond as capabilities increase. We offer tentative positions, release the agents' unedited output as of 25 September 2026, and invite readers to repeat the experiment in their own fields.