Ceo arena tests ceo agents competing in simulated markets

CEO Arena: Evaluating Long-Horizon Multi-Agent Decision-Making in Competitive Markets

Multiagent Systems

Summary

Running a business involves making many decisions over time while dealing with uncertain markets and competitors trying to outsmart each other. The authors created CEO Arena, a simulation where AI agents act as CEOs managing companies, deciding prices, marketing, and development with limited info. They tested multiple AI CEOs to see how well they perform alone and how their actions affect competitors and the market overall. The results show that even successful private gains can harm the market, revealing complex interactions among competing decision-makers.

What this means in practice

  • For business simulation designers: Simulate competitive business decisions over long timeframes to test and compare AI strategies with realistic market feedback.
  • For financial technology developers: Develop adaptive AI CEO agents that can make complex business decisions under uncertainty to improve automated financial management tools.$Commercial implications: The paper enables creating competitive AI financial decision-makers that could be sold for automated business management platforms.

Tested on simulated data.

Authors

An Yan, Yu Huo, Zhiwei Shang, Yiran Peng, Chenglin Wu

Abstract

Long-horizon competition tests agents' ability to coordinate business decisions under uncertainty and adapt to changing rival strategies. We introduce CEO Arena, a benchmark that uses matched replacement evaluation to assess operating returns alongside an agent's effects on rivals and the market. Each CEO agent is compared with a reference policy in the same company under the same economic seed, holding other agents' identities and assignments fixed while all agents adapt. In a shared eight-company market spanning 500 simulated days, CEOs make sequential decisions on pricing, procurement, marketing, research and development, and service using private company information and noisy market signals, under resource constraints and delayed feedback. We evaluate eight LLM-based CEO agents in 27 main runs and 26 robustness runs. In the main evaluation, most agents have negative mean returns, and private gains can accompany market losses. Robustness analyses suggest that aggregate patterns extend beyond the original rule-based baseline; four of the 56 directed pairs show relatively stable effects. Memory, action, and accounting traces suggest demand capture and rivals' pricing and spending responses as possible explanations. CEO Arena provides a controlled testbed for studying long-horizon agent competition, strategic interaction, and market externalities.