Large scale analysis reveals patterns in AI agent serving workloads
Semantics, Workflows, and Infrastructure: Understanding Agent Serving at Production Scale
Distributed, Parallel, and Cluster Computing
Summary
Running AI agents that use large language models at scale creates many complex requests with different actions and user inputs. The authors studied over 11 million requests on a big production system with thousands of GPUs to understand how tasks start, how workflows run, and what demands are placed on hardware. They found that request loads are very uneven, some related tasks rarely run at the same time, and some context information gets reused across tasks. These findings help reveal challenges and guide improvements for building and running large AI agent platforms.
What this means in practice
- •For ai platform engineers: Improve resource allocation and scheduling by understanding real-world AI agent request patterns and execution overlaps on large GPU clusters.
- •For cloud service providers: Design better GPU infrastructure deployments tailored to actual agent-driven workloads with skewed request volumes and context reuse.
Authors
Yihao Zheng, Jingzhe Jiang, Dejiang Zhu, Zhiyuan Tan, Yang Tian, Tao Wang, Minchen Yu
Abstract
Large language model (LLM) agents execute applications through a workflow of inference requests with tool calls and user interactions. Serving these applications at production scale requires understanding how application behavior shapes inference demand and for guiding efficient execution. Recent characterization studies provide request-level workload measurements and agent execution analysis. However, an end-to-end view connecting task initiation, workflow execution, and inference infrastructure remains unexplored. In this paper, we analyze a two-week trace of 11.7 million requests from a large-scale production platform for general-purpose agents, backed by inference infrastructure comprising over 10k GPUs. We characterize the platform at three connected levels: task-level initiation semantics, workflow-level execution patterns, and infrastructure level serving demands. Our measurements reveal workload patterns such as highly skewed request volumes across sessions, rare execution overlap among logical sibling requests, and context reuse across task boundaries. Building on these observations, we analyze deployment implications and identify open problems to guide future research on agent serving systems.