An Automatic Language Identification System for Code-Mixed English-Kannada Social Media Text

B. S. Sowmya Lakshmi, B. R. Shambhavi · 2017

The task of identifying the language of a document or word automatically is known as Language Identification (LID). With the increase in popularity of social media and smart devices, a huge number of people have come online. Majority of the user-generated data on web are code-mixed or multi-script form, where the words are represented in a non-native script. In this work, we focused on the problem of word-level LID for code-mixed data. Dataset collected contains English and Kannada code mixed sentences from social media posts. Experiments on various supervised classifiers are performed by embedding a dictionary module to handle word level code mixing.

Read the paper · More papers on PaperTik