On-device streaming improves efficiency for multimodal large language models
OmniTide: Co-Designing Algorithms and Systems for Efficient On-Device Omni-LLM Streaming
Artificial IntelligenceMachine Learning
Summary
Processing large amounts of mixed data like text, images, and more on devices like phones is hard because it uses a lot of memory and computing power. The authors found ways to smartly keep only the most important pieces of data and organize this information efficiently. This helps devices handle continuous streams of data while using less memory and running faster. Their method also keeps accuracy high and lets edge devices work with large contexts in real time.
What this means in practice
- •For mobile system developers: Create more efficient on-device AI models that process mixed data streams without running out of memory or slowing down.
- •For consumer device manufacturers: Build devices capable of real-time, long-context multimodal AI processing without relying on costly or privacy-risking cloud services.$Commercial implications: Enables new edge AI products with improved privacy and reduced operating costs for consumers and manufacturers.
Authors
Zongshang Shen, Wangsong Yin, Daliang Xu, Mengwei Xu, Xuanzhe Liu
Abstract
On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: the continuous influx of multimodal data rapidly exhausts constrained memory and compute budgets via monotonic KV cache growth. Existing sparse attention methods fall short, either incurring prohibitive online estimation latency or destroying interleaved cross-modal context, while failing to resolve physical memory fragmentation. We present OmniTide, the first algorithm-system co-design tailored for efficient on-device streaming omni-modal inference. Driven by the observation of modality-aware structural sparsity, OmniTide adopts a unit-based abstraction with two components: (1) At the algorithm level, OmniPick logically retains critical multimodal context based on unit boundaries and modality importance to preserve task accuracy; (2) At the system level, OmniPage physically partitions the cache by retention likelihood and dynamically compacts surviving sparse tokens, minimizing both memory fragmentation and data-movement overhead. Extensive evaluations across three streaming benchmarks and two consumer-device architectures show that OmniTide achieves up to $12.72\times$ kernel speedups and $2.40\times$ lower stream-loop latency. On StreamingBench, it improves accuracy by up to 18.0 percentage points over sliding-window baselines at comparable session cost. OmniPage further reduces the physical KV span by up to 26.7% relative to native logical eviction, unlocking real-time, infinite-context streaming on edge devices.