Vision language models may rewrite text based on scene context
When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context
Computer Vision and Pattern RecognitionArtificial Intelligence
Summary
Some AI models that read text in pictures can be influenced by what else is in the image, sometimes changing the text they read to match the scene better. The authors created 781 test images to see how often this happens and found it varies between models, with some rewriting text over half the time. They also showed that removing surrounding context or blurring the text changes how often models rewrite versus read literally. This means these AI models balance reading the actual letters with what the scene suggests, using context especially when the text is unclear.
What this means in practice
- •For mobile app developers: Improve text reading apps by optimizing model use of context to avoid rewriting clear text in photos.
- •For robotics engineers: Design robots that read environmental text accurately by adjusting models to weigh visual clarity over scene context.
Authors
Yuxing Cheng, Yuan Wu, Yi Chang
Abstract
Vision-language models (VLMs) can read text in natural scenes, but their predictions may be influenced by the surrounding context. When the printed text conflicts with what the scene suggests, a model may return a more plausible word instead of the shown text. We introduce SceneFaith, a benchmark of 781 generated scene images for studying this behavior. Each output is classified as Literal, Canonical, or Other, separating faithful transcription from context-consistent rewriting and ordinary recognition errors. Across 15 models from seven families, all models show rewriting on clear images, with rates ranging from 8.45\% to 58.51\%. Controlled experiments further show that surrounding context matters: removing surrounding scene information reduces rewriting and improves literal accuracy, while changing the scene around the same text patch can also change model outputs. Moreover, weakening the target text with blur increases rewriting. These results show that reliable scene-text recognition requires VLMs to balance visual character evidence with contextual information, preserving clear text while using context mainly when the visual evidence is uncertain.