Challenge
A global Document AI company needed to train OCR models for handwritten Indic scripts but had no reliable training corpus. Existing public datasets were dominated by printed text and English/Latin scripts, leaving Devanagari, Tamil, and Telugu handwriting drastically underrepresented. Handwriting varies enormously across writers, regions, and pen types - character shapes, ligatures, and conjunct consonants in Indic scripts are especially hard to generalize. Without a large, accurately annotated, multi-script handwritten dataset, the client's OCR models plateaued at unusable accuracy levels (Devanagari as low as 68%), blocking deployment for document digitization use cases across banking, government, and legal sectors.
Solution
Phase 1 Handwriting Sample Collection: eQourse recruited a diverse pool of writers across age groups, regions, and educational backgrounds to produce natural handwriting samples in Devanagari, Tamil, and Telugu, covering varied vocabulary, sentence structures, numerals, and common document formats (forms, letters, applications).
Phase 2 Multi-Level Annotation: Each image was annotated at line, word, and character level with precise bounding boxes and transcriptions. Annotators fluent in each script handled conjunct consonants, ligatures, and diacritics, with a second-pass review layer to catch transcription errors and ambiguous handwriting.
Phase 3 Data Cleaning and Delivery: The dataset was cleaned to remove illegible or duplicate samples, normalized for consistent formatting, and validated against inter-annotator agreement thresholds. Final delivery was structured in COCO JSON format with full metadata for direct ingestion into the client's OCR training pipeline.
Results
Key outcomes: Delivered 100,000+ annotated handwritten document images across 3 Indic scripts. Devanagari OCR accuracy improved from 68% to 94%, Tamil reached 91%, and Telugu reached 89%. The client was able to deploy production-grade handwritten OCR for document digitization workflows across banking, government, and legal use cases, significantly reducing manual data entry and processing time.