Papers for

mobile system developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

On-device streaming improves efficiency for multimodal large language models

OmniTide: Co-Designing Algorithms and Systems for Efficient On-Device Omni-LLM Streaming

Abstract: On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: the continuous influx of multimodal data rapidly exhausts constrained memory and compute budgets via monotonic KV cache growth. Existing sparse attention methods fall short, either incurring prohibitive online estimation latency or destroying interleaved cross-modal context, while failing to resolve physical memory fragmentation. We present OmniTide, the first algorithm-system co-design tailored for efficient on-device streaming omni-modal inference. Driven by the observation of modality-aware structural sparsity, OmniTide adopts a unit-based abstraction with two components: (1) At the algorithm level, OmniPick logically retains critical multimodal context based on unit boundaries and modality importance to preserve task accuracy; (2) At the system level, OmniPage physically partitions the cache by retention likelihood and dynamically compacts surviving sparse tokens, minimizing both memory fragmentation and data-movement overhead. Extensive evaluations across three streaming benchmarks and two consumer-device architectures show that OmniTide achieves up to $12.72\times$ kernel speedups and $2.40\times$ lower stream-loop latency. On StreamingBench, it improves accuracy by up to 18.0 percentage points over sliding-window baselines at comparable session cost. OmniPage further reduces the physical KV span by up to 26.7% relative to native logical eviction, unlocking real-time, infinite-context streaming on edge devices.

Mon 28 SeptArtificial IntelligenceMachine Learning
The gist
Processing large amounts of mixed data like text, images, and more on devices like phones is hard because it uses a lot of memory and computing power. The authors found ways to smartly keep only the most important pieces of data and organize this information efficiently. This helps devices handle continuous streams of data while using less memory and running faster. Their method also keeps accuracy high and lets edge devices work with large contexts in real time.
Open → 2609.34653v1