Decoupling semantics from vision: A framework for faithful visual-text compression evaluation

2026-08-03Computer Vision and Pattern Recognition

Computer Vision and Pattern Recognition
AI summary

The authors explain that current ways of checking how well visual-text compression (VTC) works rely too much on how well computers perform on related tasks, which doesn't really show if the original text is kept accurately. They created a new way to test VTC that separates the computer’s language understanding from the actual quality of the compression. To do this, they made a new ZeroSense Benchmark that uses tests without meaningful text connections, so results truly reflect how good the compression is. Their experiments show that good task performance doesn’t always mean good compression quality, proving their new testing method is needed.

Visual-text compressionMultimodal Large Language ModelsDeepSeek-OCRDownstream task performanceText-to-image renderingZeroSense BenchmarkSemantic correlationEvaluation frameworkLanguage priorsToken compression ratio
Authors
Yonghan Gao, Zehong Chen, Lijian Xu, Jingzhi Chen, Jingwei Guan, Xingyu Zeng
Abstract
Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image rendering. However, existing evaluation protocols heavily rely on downstream task performance. Such evaluation metrics fail to accurately measure text preservation due to the strong inherent linguistic priors of Multimodal Large Language Models (MLLMs). In this work, we introduce a new evaluation framework that decouples MLLMs' capabilities to faithfully assess VTC quality. Within this framework, we further introduce the ZeroSense Benchmark to ensure low semantic correlation of testing samples. By eliminating textual dependencies, our benchmark guarantees that the evaluation results are purely reflective of VTC quality, unaffected by the semantic inference capabilities of downstream models. Extensive experiments across multiple datasets demonstrate that VTC quality and downstream task accuracy diverge significantly, highlighting the necessity of our decoupled evaluation framework.