Machine Learning Approach for Language Identification & Transliteration

Deepak Kumar Gupta, Shubham Kumar, Asif Ekbal · 2015

In this paper, we describe the system that we developed as part of our participation to the FIRE-2014 Shared Task on Transliterated Search. We participated only for Subtask 1 that focused on labeling the query words. The entire process consists of the following components: language identification of each word in text, named entity recognition and classification (NERC) and transliteration of Indian language words written in non-native scripts to the corresponding native Indic scripts. The proposed methods of language identification and NERC are based on supervised approaches, where we use several machine learning algorithms. Our transliteration framework is based on modified joint source channel model. Experiments on benchmark setup show that we achieve quite encouraging performance for both the pairs of languages, viz. Bangla-English and Hindi-English. It is also to be noted that we did not make use of heavy domain-specific resources and/or tools, and therefore this can be easily adapted to the other domains and/or languages.

Read the paper · More papers on PaperTik