A language identification system for code-mixed English-Manipuri Social Media text
Priyadarshini Lamabam, Kunal Chakma · 2016
In Social Media, code mixing and code switching take place where users often communicate in two or more languages. The state-of-the-art techniques fail when language identification is done on such code mixed informal texts due to the presence of lexical borrowings, creative spellings and phonetic typing. Therefore, automatic language identification for the code mixed social media texts has become a challenging task in the field of Natural Language Processing(NLP). A dataset was selected containing Twitter and Facebook posts that exhibit code-mixing between English and Manipuri. Some word-level language identification experiments are performed using this dataset such as Trigram-based and Conditional Random Field (CRF)-based models. We find that the CRF-based models have given better accuracies in identifying the languages.