Papers for

media analysis platform developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Efficient serving system balances GPU use for multimodal language models

EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models

Abstract: Disaggregating the two stages, Prefill and Decode, onto separate GPU pools is now a standard optimization for (text-only) LLM serving. However, multimodal LLMs (MLLMs), which add a third phase, Encode, pose new challenges for resource allocation. Encode turns images, video, or audio into embeddings that the language model can consume, yielding a three-stage Encode-Prefill-Decode (EPD) pipeline. Existing frameworks offer only partial answers: text-only PD systems lack Encode, while EPD frameworks expose it as a separate service without regulating downstream request flow. The pipeline also carries a structural resource imbalance: every request enters through Encode before downstream work can begin, yet per-request execution leaves the encode GPU severely underutilized even at high loads, starving the downstream Prefill and Decode workers. Addressing this, we reposition Encode as the control point of the EPD pipeline, exposing three tightly coupled dimensions: when work enters downstream, where prefill executes, and how the GPU is shared. We instantiate this in EAServe across two co-designed layers. Its runtime manages load-adaptive micro-batching, rate-controlled partial offload to a co-resident prefill worker, and dynamic SM partitioning for predictable co-location. The configuration layer, Hybrid Auto Selection (HAS), navigates the joint space of GPU allocation, encode batch size, and offload ratio by pruning unbalanced allocations with per-stage capacity profiling and refining the remainder through TPE-based Bayesian optimization. Evaluated on three MLLM architectures spanning image, video, and audio, EAServe delivers up to 4.3x and 1.7x higher goodput than NVIDIA Dynamo and vLLM, respectively, under identical SLO constraints, sustains more balanced and higher GPU utilization across the EPD pipeline, and reaches near-optimal configurations faster than baseline search methods.

Fri 25 SeptDistributed, Parallel, and Cluster ComputingMachine LearningPerformance
The gist
Multimodal large language models process images, video, or audio along with text, requiring an extra Encode step that turns media into embeddings. Existing serving systems struggle to allocate GPU resources well among Encode, Prefill, and Decode steps, leading to wasted capacity. The authors propose EAServe, a system that controls when and how work flows through these three steps and allocates GPUs dynamically to keep all stages busy. This approach improves throughput and GPU utilization compared to prior methods under the same delay goals. EAServe was tested on models handling image, video, and audio inputs, showing faster and more balanced performance.
Open → 2609.31551v1