Marathoner enables AI to work for hours on complex tasks autonomously

Marathoner: Ultra-Long-Horizon Autonomous Intelligence

Computer Vision and Pattern Recognition

Summary

Humans can work for a long time to finish big, complicated jobs, but computers usually can’t do that without help. The authors built Marathoner, an AI model that can independently handle very long and difficult tasks lasting many hours with thousands of steps. They taught it using real large coding projects from GitHub and combined lots of smaller tasks into bigger ones. Marathoner got better by training on trial attempts and getting rewards especially for doing well in later parts of tasks. It can now work for over 10 hours and make a thousand+ calls to tools while solving hard problems.

What this means in practice

  • For software development teams: Automate the generation and integration of large-scale code changes by running Marathoner on complex multi-step programming tasks.
  • For automation engineers: Deploy Marathoner to manage and execute extended sequences of tasks requiring multiple tool interactions over long periods.

Authors

Zhang Ruiyang, Ou Jinpeng, Xie Yifan, Zhou Jingang, Pan Lirui, Guo Qingpei, Zheng Zhedong

Abstract

Humans naturally possess the ability to work persistently toward long-term goals. Given a challenging task, humans can continuously work for months or even years to accomplish a specific objective. In this paper, we propose Marathoner, an autonomous agentic model possessing the ability of ultra-long-horizon execution. Specifically, we propose a comprehensive post-training pipeline to instill this critical capability into base model. For Ultra-Long-Horizon Task Synthesis, we leverage major release PRs containing 1000+ lines of new code from diverse GitHub repositories as the primary source for synthesizing challenging task-level data. Additionally, we introduce Multi-Task Chaining, which chains multiple generated tasks into a single more challenging task, enabling the synthesis of tasks with frontier-level difficulty. For rejection sampling finetuning, we combine strong teacher model with diverse harnesses to generate trajectories on our synthesized tasks and conduct supervised finetuning on base model with rejection sampled trajectories. For reinforcement learning, cold-started model performs real-world execution through harnesses in independent sandboxes during rollout process, effectively facilitating the acquisition of genuine ultra-long-horizon execution capability. We further propose a novel reward strategy, Later Stage Bonus Reward, which explicitly encourages model to perform meaningful maneuvers during later stages of execution. Through extensive evaluation on 5 benchmarks containing ultra-long-horizon tasks, Marathoner achieves consistent and substantial performance improvements over base model and even surpasses performance of strong proprietary model. Further analysis shows that Marathoner can consistently work for 10+ hours and conduct 1000+ tool calls on highly challenging tasks.