Handwritten text recognition improves by focusing on high-variance pixels
Handwritten Text Recognition Lives in the High-Pixel Variance Subspace
Computer Vision and Pattern RecognitionMachine Learning
Summary
Handwritten text recognition systems usually learn by either reconstructing images or contrasting differences between them. This paper shows that for handwriting, the important details lie in parts of the image that vary the most, so methods that focus on reconstructing these high-variance areas work better than contrast-based methods. The authors tested different methods across varied handwriting datasets and found that systems concentrating on these high-variance pixels were better at recognizing text, especially when there were fewer labels available. Using these findings, their system even competes well with fully trained supervised methods.
What this means in practice
- •For handwriting recognition developers: Use pixel reconstruction focusing on high-variance pixels to improve handwriting recognition accuracy even with limited labeled data.
- •For document digitization teams: Deploy pretrained pixel-grounded handwriting models to transcribe diverse handwriting styles across multiple languages with improved error rates.
Authors
Carlos Garrido-Munoz, Jorge Calvo-Zaragoza
Abstract
In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel space: for HTR, it is concentrated in high-variance directions and largely absent from low-variance ones. This predicts that objectives preserving high-variance pixel content will transfer best. We test six SSL methods from three families (pixel-grounded MIM, JEPA, and contrastive) under matched encoder, data, and evaluation protocols on six handwriting benchmarks across five languages. With full labels, pixel-groundrounded SSL achieves the lowest CER on every benchmark and both frozen probes, exposes per-position character information that other families recover only through the readout, and is the only family to benefit from pretraining on real handwriting. Pixel-grounded representations are also more label efficient. Across datasets, encoder alignment with the high-variance pixel subspace predicts CER within every method. With a pretrained LLM decoder, a frozen pixel-grounded encoder is competitive with fully fine-tuned supervised baselines; full fine-tuning achieves the lowest mean CER and ranks first or second on every benchmark. These results show that the value of pixel reconstruction depends on where discriminative signal lies in the input.