Papers for

handwriting recognition developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Handwritten text recognition improves by focusing on high-variance pixels

Handwritten Text Recognition Lives in the High-Pixel Variance Subspace

Abstract: In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel space: for HTR, it is concentrated in high-variance directions and largely absent from low-variance ones. This predicts that objectives preserving high-variance pixel content will transfer best. We test six SSL methods from three families (pixel-grounded MIM, JEPA, and contrastive) under matched encoder, data, and evaluation protocols on six handwriting benchmarks across five languages. With full labels, pixel-groundrounded SSL achieves the lowest CER on every benchmark and both frozen probes, exposes per-position character information that other families recover only through the readout, and is the only family to benefit from pretraining on real handwriting. Pixel-grounded representations are also more label efficient. Across datasets, encoder alignment with the high-variance pixel subspace predicts CER within every method. With a pretrained LLM decoder, a frozen pixel-grounded encoder is competitive with fully fine-tuned supervised baselines; full fine-tuning achieves the lowest mean CER and ranks first or second on every benchmark. These results show that the value of pixel reconstruction depends on where discriminative signal lies in the input.

Mon 28 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Handwritten text recognition systems usually learn by either reconstructing images or contrasting differences between them. This paper shows that for handwriting, the important details lie in parts of the image that vary the most, so methods that focus on reconstructing these high-variance areas work better than contrast-based methods. The authors tested different methods across varied handwriting datasets and found that systems concentrating on these high-variance pixels were better at recognizing text, especially when there were fewer labels available. Using these findings, their system even competes well with fully trained supervised methods.
Open → 2609.35473v1

Handwritten devanagari recognition needs fewer transcriptions with pretraining

Measuring Annotation Efficiency for Handwritten Devanagari Recognition: Sample-Complexity Curves for Four Pretraining Regimes

Abstract: To train handwritten text recognition systems we need word images and their corresponding transcriptions, and these transcriptions are produced manually. For a script that can be read by only a small number of specialists, this manual transcription is a limitation, because the trained models are supposed to save the time of those same specialists. A relevant question therefore arises: how many transcriptions are needed before a recogniser becomes useful, and how much of that cost can pretraining remove? In this study the answer is measured directly for handwritten Devanagari. We keep the recogniser, optimiser and evaluation protocol the same and change only the number of real transcribed words used for fine-tuning across nine budgets from 10 to 4,000 and four initialisation regimes, with six seeds at every point. The resulting curves are then converted into annotation-equivalent terms. A CER of 0.50 is reached by supervised synthetic pretraining using only 81 transcribed words, whereas random initialisation requires 355, which gives a label multiplier of 4.40 [3.56, 4.99]. There is a zero-shot reference point as well: with no real transcribed words at all, this pretraining is worth about 136 of them. This advantage gets smaller as the target accuracy improves, and at the most demanding target we measure, it cannot be distinguished from no saving at all. A fourth arm in which only the encoder is transferred separates the effect of the pretraining method from that of transfer scope, and masked image modelling is observed to transfer negatively over a bounded range of budgets. We emphasise that the scarcity in this study is constructed by subsampling a large corpus.

Tue 15 SeptComputer Vision and Pattern RecognitionMachine Learning
The gist
Training systems to recognize handwritten Devanagari text usually requires many manual transcriptions, which is hard when only a few experts can read the script. The authors tested how many transcriptions are really needed to get a usable model and how pretraining on synthetic data helps. They found that pretraining can reduce the needed transcriptions by over four times to reach a moderate error level. However, this benefit shrinks as higher accuracy is targeted, and some pretraining methods can even hurt performance in certain conditions.
Open → 2609.16859v1