Back to Research papers
Research paper index

Koshur Pixel: a large-scale synthetic ocr dataset for kashmiri

Haq Nawaz Malik, Faizan Iqbal, Nahfid Nissar

arXiv:2606.23144Published June 22, 20260 citations
  • cs.CV
  • cs.CL

Abstract

Optical Character Recognition (OCR) for low-resource languages is often constrained by the lack of annotated training data and the complexity of script-specific rendering. Kashmiri, written primarily in the Perso-Arabic Nastaliq script, presents additional challenges due to contextual glyph shaping, dense ligatures, and orthographic variability. We introduce Koshur Pixel, the first large-scale synthetic OCR dataset for Kashmiri, comprising 613,078 image-text pairs generated from the KS-PRET-5M corpus using the SynthOCR-Gen framework. The dataset spans multiple fonts and textual granularities, ranging from individual words to full-page documents, and incorporates more than 25 augmentation strategies that emulate real-world document degradations. Koshur Pixel provides a scalable and cost-effective alternative to manual annotation, establishing a foundational resource for training OCR systems, digitizing Kashmiri textual heritage, and advancing language technologies for a severely under-resourced language.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.