Identification of Multilingual Words Using Profile Based Features

Kapil Kumar Kaswan · 2014

In a multi Language environment, majority of the documents may contain text information printed in more than one script/language. For automatic processing of such documents through Optical Character Recognition, it is necessary to identify different Language regions of the document. In this dissertation, it is proposed to develop a model to identify the Language type of trilingual words printed in Punjabi, Hindi and English Languages. The distinct characteristic features of Punjabi, Hindi and English Languages are thoroughly studied from the nature of the top bottom and middle profiles. The proposed model is applied on the words present in all the three languages. Experimentation conducted involved 200 text words for testing. The results are encouraging and prove the efficiency of the proposed model. The average success rate is found to be 94.50% for data set constructed from scanned document images and created data set.

Read the paper · More papers on PaperTik