Efficient serving system balances GPU use for multimodal language models

EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models

Distributed, Parallel, and Cluster ComputingMachine LearningPerformance

Summary

Multimodal large language models process images, video, or audio along with text, requiring an extra Encode step that turns media into embeddings. Existing serving systems struggle to allocate GPU resources well among Encode, Prefill, and Decode steps, leading to wasted capacity. The authors propose EAServe, a system that controls when and how work flows through these three steps and allocates GPUs dynamically to keep all stages busy. This approach improves throughput and GPU utilization compared to prior methods under the same delay goals. EAServe was tested on models handling image, video, and audio inputs, showing faster and more balanced performance.

What this means in practice

Authors

Kunxiong Zhu, Zhihao Shu, Hangyu Zheng, Minghai Qin, Miao Yin, Gagan Agrawal, Wei Niu

Abstract

Disaggregating the two stages, Prefill and Decode, onto separate GPU pools is now a standard optimization for (text-only) LLM serving. However, multimodal LLMs (MLLMs), which add a third phase, Encode, pose new challenges for resource allocation. Encode turns images, video, or audio into embeddings that the language model can consume, yielding a three-stage Encode-Prefill-Decode (EPD) pipeline. Existing frameworks offer only partial answers: text-only PD systems lack Encode, while EPD frameworks expose it as a separate service without regulating downstream request flow. The pipeline also carries a structural resource imbalance: every request enters through Encode before downstream work can begin, yet per-request execution leaves the encode GPU severely underutilized even at high loads, starving the downstream Prefill and Decode workers. Addressing this, we reposition Encode as the control point of the EPD pipeline, exposing three tightly coupled dimensions: when work enters downstream, where prefill executes, and how the GPU is shared. We instantiate this in EAServe across two co-designed layers. Its runtime manages load-adaptive micro-batching, rate-controlled partial offload to a co-resident prefill worker, and dynamic SM partitioning for predictable co-location. The configuration layer, Hybrid Auto Selection (HAS), navigates the joint space of GPU allocation, encode batch size, and offload ratio by pruning unbalanced allocations with per-stage capacity profiling and refining the remainder through TPE-based Bayesian optimization. Evaluated on three MLLM architectures spanning image, video, and audio, EAServe delivers up to 4.3x and 1.7x higher goodput than NVIDIA Dynamo and vLLM, respectively, under identical SLO constraints, sustains more balanced and higher GPU utilization across the EPD pipeline, and reaches near-optimal configurations faster than baseline search methods.