Multilingual Language Identification: ALTW 2010 Shared Task Data

Timothy J. Baldwin, Marco Lui · 2010

While there has traditionally been strong interest in the task of monolingual language identification, research on multilingual language identification is underrepresented in the literature, partly due to a lack of standardised datasets. This paper describes an artificially-generated dataset for multilingual language identification, as used in the 2010 Australasian Language Technology Workshop shared task.

Read the paper · More papers on PaperTik