Language Detection of Text Document Image
N. Jayanthi, Harsha Harsha, Naman Jain, Ishpreet Singh Dhingra · 2020
Language detection is an important feature for developing a Multilingual OCR. Lots of research has been done on Languages such as English, French, Spanish but very limited work has been performed on Indian languages such as Hindi, Tamil etc. due to their complex nature. In this paper we are presenting a method to detect the language of a text document image. The language chosen are such that they are of different scripts. The process of language detection involves segmentation of individual characters from the document and then labelling them to a particular language. A Convolutional Neural Network is trained on characters of different languages. A text document is classified on the basis of what language its individual characters belong to. This is achieved by first segmenting individual characters of a document and then classifying each of them separately.