Automatic Token and Turn Level Language Identification for Code-Switched Text Dialog: An Analysis Across Language Pairs and Corpora

Vikram Ramanarayanan, Robert Pugh · 2018

We examine the efficacy of various feature-learner combinations for language identification in different types of text-based code-switched interactionshuman-human dialog, human-machine dialog, as well as monolog -at both the token and turn levels.In order to examine the generalization of such methods across language pairs and datasets, we analyze ten different datasets of code-switched text.We extract a variety of character-and word-based text features and pass them into multiple learners, including conditional random fields, logistic regressors, and recurrent neural networks.We further examine the efficacy of character-level embedding and GloVe features in improving performance and observe that our best-performing text system significantly outperforms the majority vote baseline across language pairs and datasets.

Read the paper · More papers on PaperTik