Football management agents tested for long term decision impact

FromPitch2Board: Benchmarking LLM Agents in Long-Horizon Football Management

Artificial Intelligence

Summary

Football management AIs are tested to see how different factors affect their performance over a long time. The researchers created FromPitch2Board, a precise way to measure if AI skills come from the model itself or from how much control it has. They found that as the AI’s responsibilities grow, it sometimes makes fewer decisions, showing where its strengths and weaknesses lie. Different model setups lead to changes in performance that a single overall score can't explain.

What this means in practice

  • For game ai developers: Compare AI coaching strategies to find which design choices improve long-term football management tasks.
  • For sports simulation platform teams: Use a reliable benchmark to evaluate simulation realism and agent behavior changes under different control levels in football management games.

Authors

Peiyu Zang

Abstract

Long-horizon agent benchmarks typically report how far an agent progresses, but do not identify whether its performance comes from the foundation model, scaffold, responsibility scope, match-control granularity, or horizon. We introduce FromPitch2Board, a deterministic football-management benchmark that studies five configurable factors through controlled comparisons on a single simulator, using paired seeds and a frozen calibration. We evaluate four foundation models and four agent scaffolds. In the Model Track, Coach points Z-scores span 0.19, while Manager points Z-scores span 0.68, with GPT-5.6 showing a sharp rise in passivity under responsibility expansion. Its responsibility ladder rises from 46.1 to 58.1 points with recruitment, then falls to 46.8 under full management, localizing the regression to the final responsibility boundary. Across that boundary, its skipped-decision rate rises from 1.1% to 57.9%. Within the Flash-Pro pair crossed across every scaffold, scaffold choice changes Manager points Z-scores by up to 0.48 relative to the fixed stateless scaffold. The 3Y cohort shows a directional reversal in mean ranking between years one and three, while a selected Claude Code+Pro configuration peaks in year three and remains below that peak, showing that responsibility scope and horizon expose behavior changes that a single headline score conceals.