Dynamic resolution routing improves multimodal model efficiency

Resolution as a First-Class Decision: Task-Conditioned Routing for Efficient Multimodal Large Language Models

Computer Vision and Pattern Recognition

Summary

Multimodal large language models process both text and images, but handling high-resolution images uses a lot of computing power. The researchers found that always using the same image resolution is inefficient because different tasks might need different levels of detail. They designed a system that decides the best image resolution based on the task using both the image and text input. This method reduces the computing work and speeds up the model without hurting its performance.

What this means in practice

  • For ai system engineers: Deploy more efficient multimodal models by dynamically adjusting image resolution per task to reduce computational requirements and latency.
  • For mobile app developers: Integrate resource-saving image processing in apps combining text and visuals, improving speed and battery life on devices.

Authors

Zhiqiang Xia, Yang Li, Xinyuan Zhang, Yuchen Liu, Haoyu Lu, Jiaming Xu, Runyu Shi, Ying Huang

Abstract

The inference efficiency of Multimodal Large Language Models (MLLMs) is severely constrained by massive visual token sequences induced by high-resolution inputs, with computational cost scaling quadratically. Existing approaches primarily focus on downstream token compression, while overlooking a fundamental upstream inefficiency: input resolution is treated as a static, task-agnostic hyperparameter. We propose Task-Conditioned Resolution Routing (TCRR), which formulates visual compression as a task-conditioned decision and employs a lightweight cross-modal router that conditions backbone visual representations on textual semantics via feature-wise modulation and cross-attention to predict the minimal sufficient compression level per query. To support this, we curate a dataset of 500k samples across 12 task categories, labeled via a teacher-oracle pipeline to approximate Pareto-optimal compression scales. Extensive experiments across diverse architectures show that TCRR achieves a superior efficiency frontier, specifically reducing visual FLOPs by 40.9% and latency by 53.7% on Qwen3-VL-8B while preserving competitive performance. Further analysis of scaling behavior confirms that dynamically routing visual compression enables optimal resource allocation without modifying the MLLM backbone.