A Language Modeling Approach to Identifying Code-Switched Sentences and Words
Liang Yu, Wei-Cheng He, Wei-Nan Chien · 2012
Globalization and multilingualism contribute to code-switching – the phenomenon in which speakers produce utterances containing words or expressions from a second language. Processing code-switched sentences is a significant challenge for multilingual intelligent systems. This study proposes a language modeling approach to the problem of codeswitching language processing, dividing the problem into two subtasks: the detection of code-switched sentences and the identification of code-switched words in sentences. A codeswitched sentence is detected on the basis of whether it contains words or phrases from another language. Once the code-switched sentences are identified, the positions of the code-switched words in the sentences are then identified. Experimental results on Mandarin-Taiwanese code-switching sentences show that the language modeling approach achieved a 79.52 % F-measure and an accuracy of 80.23% for detecting code-switched sentences, and a 51.20 % F-measure for the identification of code-switched words. 1