A Neural Model for Language Identification in Code-Switched Tweets

Aaron Jaech, George Mulcaire, Mari Ostendorf, Noah A. Smith · 2016

Language identification systems suffer when working with short texts or in domains with unconventional spelling, such as Twitter or other social media.These challenges are explored in a shared task for Language Identification in Code-Switched Data (LICS 2016).We apply a hierarchical neural model to this task, learning character and contextualized word-level representations to make word-level language predictions.This approach performs well on both the 2014 and 2016 versions of the shared task.

Read the paper · More papers on PaperTik