Compile-time settings optimize waiting and batch costs in llm serving
Tool Waiting and Re-arrival in Compile-Time-Static LLM Serving: Cost Mechanisms and Configuration Selection
Performance
Summary
Large language models often use external tools during their work, which makes them pause and then continue afterward. The paper studies how fixed settings decided before running—like batch sizes and memory slots—affect the cost of these pauses and resumptions in specialized hardware. The authors created a simulator to test over two thousand different configurations and found that changing these settings can reduce the waiting cost by up to about 10%. They explain why some adjustments help and why simply looking at batch padding isn't enough to estimate efficiency.
What this means in practice
- •For ai infrastructure engineers: Choose fixed serving configurations that reduce tool wait overhead and improve cost efficiency on specialized hardware.
- •For cloud service operators: Optimize language model serving setups to better handle external tool calls and reduce inference runtime costs.
Authors
Dongkyeom Jang, In-Nea Wang, Junho Jeong
Abstract
In agentic LLM services, a session calls an external tool, waits for it, and re-arrives to continue inference. Statically compiled NPU serving can fix the batch bucket set, the maximum batch size, and the number of KV cache slots at compile time. We define such an environment as a compile-time-static serving substrate and analyze the execution-time cost that tool waiting and re-arrival incur in it. On a single LLM instance, we run synthetic workloads following a measured tool waiting time distribution and compare, on the same inputs, a baseline configuration with settings {1, 2, 4, 8}, 8, and 8 against configurations that change some of them. Because re-arrival times differ across configurations, we build a simulator that replays request processing in time order, select the candidate with the lowest predicted cost among 2,077 configurations, and validate it on new inputs. We identify three mechanisms: discrete batch alignment, KV cache survival, and prefill interference. At a concurrency of 6, absent from the bucket set, tool waiting lowered the padding ratio (0.235 to 0.120) yet increased decode execution time 1.51-fold, so padding alone did not indicate cost. On new inputs at a concurrency of 8, where the baseline reused KV in 9 of 24 re-arrivals, enlarging the maximum batch size alone cut execution cost by 8.25%, and the selected configuration, which also adjusted the bucket set, by 9.72%. Where 17 of 18 re-arrivals were already reused, the effect was 0.59%. Compile-time configurations should thus be selected by diagnosing KV reuse loss and the resulting change in execution.