Exploring Advancements in Multilingual OCR Systems for Enhanced Document Analysis and Text Recognition

Kinjal Patel, Deepak Parashar, Nilesh Bhaskarrao Bahadure, Bhoomi Shah, Rohit Kumar, Jagdish Chandra Patni · 2025

Optical Character Recognition (OCR) systems use robust software for searching words from scanned multilingual Indian documents. Manually searching such documents is tedious and time- consuming. These documents suffer from their improper layout, and low print quality, and contain intermixed texts (Machine-printed and handwritten). OCR is used to detect text from images if the text is not visible then it detects the actual text and gives the visible text. The system improves text recognition accuracy and it takes less time to identify the original text. The system uses many algorithms or methods to perform these tasks like Convolutional Neural Network (CNN), Byte Pair Encoding (BPE), and Language model (LM). It gives experimental results that demonstrate significant advancement in text recognition performance and scalability. It offers a comprehensive solution for multilingual OCR tasks.

Read the paper · More papers on PaperTik