Analysis of the key components of segmentation-free bilingual OCR for mobile phones

Deepak Kumar, A. G. Ramakrishnan · 2022 IEEE 19th India Council International Conference (INDICON) · 2022

The recognition of text from camera-captured images using mobile phones has tremendous applications. The research work to drive mobile phone-based applications is much needed and ubiquitous. One of the applications is the transliteration of text in the image from one language into another. We need to recognize the text using an OCR engine and then perform transliteration. Here, we focus on the problems encountered while developing an OCR engine to recognize bilingual (Kannada and English) text from camera-captured images. A number of components are involved in building a bilingual OCR engine. We need a large corpus of real-world images to evaluate the OCR engine on camera-captured images. We need a neural network model that can handle text in two different languages without hassle. We need the model to run on mobile phones and recognize the text in the image. In this work, we analyze the challenges involved in achieving high performance. Still, there is scope for improvement in recognizing out-of-vocabulary words, which were not part of training the model.

Read the paper · More papers on PaperTik