Evaluating the Performance of Automated Part-of-Speech Taggers on an L2 Corpus

Craig Hagerman, クレッグ ヘガマン · Institutional Repositories DataBase (IRDB) · 2011

Automated Part-of-Speech (POS) tagging is commonly on corpora in order to allow for the systematic study. POS tagging is also a fundamental stage in most natural language processing (NLP) tasks. Although there is a long history of research into automated POS tagging in the field of NLP, the vast majority of the research has been on first language texts. Increasingly second language learner corpora are being compiled. As well, increasing use of English as a second language makes the processing of non-native English texts increasingly likely for NLP applications. However, there is very little research into how second language texts affect the performance of automated POS taggers. This paper describes a study which (1) compares the performance of three taggers on native and second language texts and (2) identifies which POS tagger has the highest level of accuracy when faced with second language writing.

Read the paper · More papers on PaperTik