A Rule-based Semi-automated OCR Postprocessing Method for Aligning Multi-language Transcripts with Multi-column Text
Jarjish Rahaman, Ankita Jain, Kousik Sarathy Sridharan, Mohan Raghavan · 2023
Optical character recognition (OCR) is essential for converting physical documents into digital format, but it faces challenges with complex documents containing multi-language transcripts and multi-column text. This paper proposes a rule-based, semi-automated OCR postprocessing method to address text alignment issues and enhance the quality and readability of recognized text in such scenarios. The following approach combines techniques such as max gap detection for page separation; bounding box sorting and merging for text alignment; header and subheader extraction; and column-wise text alignment. Results demonstrate a mean alignment accuracy of 91.5%, with a standard deviation of 1.2%. The qualitative analysis further confirms the effectiveness of the proposed approach. Limitations and future research directions are discussed, emphasizing the potential for broader applicability, integration of advanced techniques, and evaluation on diverse datasets. Overall, the postprocessing method significantly improves text alignment accuracy and has implications for document digitization and information retrieval.