Word Level Language Identification in Assamese-Bengali-Hindi-English Code-Mixed Social Media Text
Neelakshi Sarma, Sanasam Ranbir Singh, Diganta Goswami · 2018
The content posted over social media platforms today are characterized by code-mixing, phonetic typing, lexical borrowing and neologism making word level language identification an important pre-requisite for various natural language processing applications. Existing studies in word level language identification are not applicable for low resource languages due to the lack of essential tools and resources like dictionaries, annotated resources, transliteration tools etc. In this paper, we address the problem of word level language identification in a highly multilingual environment for low resource languages. We investigate the performance of different word level language identification frameworks over a corpus of transliterated Assamese-Bengali-Hindi-English messages collected from Facebook social media platform. From various experimental observations, it is evident that global semantic similarities help in identifying borrowed words, and local contextual similarity helps in resolving words that are valid in multiple languages.