Papers for

sports simulation platform teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Football management agents tested for long term decision impact

FromPitch2Board: Benchmarking LLM Agents in Long-Horizon Football Management

Abstract: Long-horizon agent benchmarks typically report how far an agent progresses, but do not identify whether its performance comes from the foundation model, scaffold, responsibility scope, match-control granularity, or horizon. We introduce FromPitch2Board, a deterministic football-management benchmark that studies five configurable factors through controlled comparisons on a single simulator, using paired seeds and a frozen calibration. We evaluate four foundation models and four agent scaffolds. In the Model Track, Coach points Z-scores span 0.19, while Manager points Z-scores span 0.68, with GPT-5.6 showing a sharp rise in passivity under responsibility expansion. Its responsibility ladder rises from 46.1 to 58.1 points with recruitment, then falls to 46.8 under full management, localizing the regression to the final responsibility boundary. Across that boundary, its skipped-decision rate rises from 1.1% to 57.9%. Within the Flash-Pro pair crossed across every scaffold, scaffold choice changes Manager points Z-scores by up to 0.48 relative to the fixed stateless scaffold. The 3Y cohort shows a directional reversal in mean ranking between years one and three, while a selected Claude Code+Pro configuration peaks in year three and remains below that peak, showing that responsibility scope and horizon expose behavior changes that a single headline score conceals.

Mon 28 SeptArtificial Intelligence
The gist
Football management AIs are tested to see how different factors affect their performance over a long time. The researchers created FromPitch2Board, a precise way to measure if AI skills come from the model itself or from how much control it has. They found that as the AI’s responsibilities grow, it sometimes makes fewer decisions, showing where its strengths and weaknesses lie. Different model setups lead to changes in performance that a single overall score can't explain.
Open → 2609.34710v1