N-gram based algorithm for distinguishing between Hindi and Sanskrit texts

C Sreejith, M. Indu, P. C. Reghu Raj · 2013 Fourth International Conference on Computing, Communications and Networking Technologies (ICCCNT) · 2013

Language Identification (LI) is the process of determining the natural language in which the given content is written. It is an important preprocessing step in many tasks of Natural Language Processing (NLP). In a multilingual society like India, automatic language identification has a wider scope, since it would be a vital step in bridging the digital divide between the Indian masses and others. In this paper, we present an N-gram based method of language identification for documents written in Hindi and Sanskrit, which have the same script and the results are shown. The technique can also be applied to other pairs of Indian languages sharing common scripts.

Read the paper · More papers on PaperTik