Identifying Languages at the Word Level in Code-Mixed Indian Social Media Text
Amitava Das, Björn Gambäck · BIBSYS Brage (BIBSYS (Norway)) · 2014
Language identification at the document level has been considered an almost solved problem in some application areas, but language detectors fail in the social media context due to phenomena such as utter-ance internal code-switching, lexical bor-rowings, and phonetic typing; all imply-ing that language identification in social media has to be carried out at the word level. The paper reports a study to detect language boundaries at the word level in chat message corpora in mixed English-Bengali and English-Hindi. We introduce a code-mixing index to evaluate the level of blending in the corpora and describe the performance of a system developed to sep-arate multiple languages. 1