Text aware single step method improves text image resolution

TOLA: Text-aware One-Step Latent Adaptation for Diffusion-based Text Image Super-Resolution

Computer Vision and Pattern RecognitionArtificial Intelligence

Summary

Text image super-resolution tries to make blurry or damaged text in images clear and readable again. Previous methods took many steps and were slow, sometimes making errors worse by repeating wrong guesses. The authors propose a new method called TOLA that works in just one step and avoids repeating mistakes by carefully judging which text guesses to trust. Their method also fixes small details in the text strokes to create clearer letters. Tests show TOLA works better and faster than previous approaches.

What this means in practice

  • For mobile app developers: Improve OCR outputs and text clarity in apps that scan or photograph text under poor conditions using a faster single-step enhancement method.
  • For document management teams: Enhance readability and accuracy of scanned or degraded text images in digital archives with a robust correction technique that reduces error amplification.

Authors

Yike Xu, Yue Shi, Yong Guo, Jiezhang Cao

Abstract

Text image super-resolution (TSR) aims to recover visually faithful and readable text under unknown degradations. Existing diffusion-based methods typically rely on multi-step prediction of either the high-resolution image or its text prior, resulting in prohibitive computational cost and inference latency. More critically, an erroneous text prior may be repeatedly injected into the denoising process, causing image and text predictions to reinforce each other and progressively amplify an early recognition error into a sharp yet semantically incorrect character. To address these limitations, we propose TOLA, a Text-aware One-step Latent Adaptation framework without iterative image-text diffusion. TOLA consists of two key modules. First, a confidence-weighted text conditioning module constructs the semantic condition only once and suppresses unreliable OCR predictions before they contaminate image reconstruction. Second, a lightweight latent residual correction module explicitly estimates and corrects the structured residual errors to recover missing or distorted stroke details. Extensive experiments demonstrate our state-of-the-art performance across all evaluation metrics on both CTR-TSR-Test ($\times 4$) and RealCE-200 benchmarks. It is worth noting that our TOLA consistently surpasses existing diffusion-based TSR methods by at least 2.72 dB in PSNR on CTR-TSR-Test.