Historical persian manuscript dataset improves word spotting accuracy

HPMD: A Historical Persian Manuscript Dataset for Word Spotting with Line-Level Annotation

Computer Vision and Pattern Recognition

Summary

Searching old Persian handwritten books for specific words is hard because there are no good public datasets and marking words manually is expensive. The authors created a new large dataset of Persian manuscripts with detailed line and text labels. They also developed a method that finds lines and spots words without needing word-level marking during training, improving search accuracy. This approach handles visually similar Persian letters better, making it easier to locate words in the text.

What this means in practice

  • For digital archive managers: Improve search tools to quickly find specific words in large collections of historical Persian manuscripts using line-level annotations instead of costly word-level labeling.
  • For optical character recognition developers: Develop OCR systems that can better locate and identify words in handwritten Persian text by using frame-level character posterior matching instead of exact decoded text.

Tested on one dataset.

Authors

Saeid Firouzi Daghigh, Majid Iranpour Mobarakeh

Abstract

Large collections of historical Persian manuscripts have been digitized, but searching them is still slow and mostly manual. Historians usually want to find where a specific name, date, event, or topic appears, which is a word spotting problem. Progress on this task is limited by two things. First, there is almost no public dataset of historical Persian handwriting; the only notable resource, OpenITI MAKHZAN, contains a relatively small Persian portion. Second, word spotting models usually need word-level bounding boxes, which are very expensive to annotate. In this paper we introduce a new dataset of 223 pages, 3,678 lines, 37,631 words, and 130,630 characters, collected from diverse historical Persian books of poetry and prose and annotated at the region, line, and text level. We also propose a baseline that is trained only with line-level annotations but returns word-level locations. A fine-tuned line detector finds text lines, and a fine-tuned CRNN recognizer trained with CTC produces a frame-by-character posterior matrix for each line. Instead of decoding the most probable character at each frame, the query is scored directly against this matrix, so visually similar characters in Persian such as be and pe no longer cause hard failures. The frame alignment also gives the horizontal position of the word inside the line. On the test set, the fine-tuned line detector reaches an F1 of 0.892, and posterior-based search raises the word spotting F1 from 0.487 to 0.558 compared with exact matching on the decoded text, with the decision threshold selected on a held-out validation set. A PHOC attribute-embedding baseline that additionally receives oracle word boundaries at test time reaches an F1 of 0.449, below the proposed method. We also report a distributional analysis of the dataset, a taxonomy of retrieval errors, and a per-conditionbreakdown of performance.