The Effect of Learner Corpus Size in Grammatical Error Correction of ESL Writings

Tomoya Mizumoto, Yuta Hayashibe, Mamoru Komachi, Masaaki Nagata, Yūji Matsumoto · 2012

English as a Second Language (ESL) learners ’ writings contain various grammatical errors. Pre-vious research on automatic error correction for ESL learners ’ grammatical errors deals with re-stricted types of learners ’ errors. Some types of errors can be corrected by rules using heuristics, while others are difficult to correct without statistical models using native corpora and/or learner corpora. Since adding error annotation to learners ’ text is time-consuming, it was not until recently that large scale learner corpora became publicly available. However, little is known about the ef-fect of learner corpus size in ESL grammatical error correction. Thus, in this paper, we investigate the effect of learner corpus size on various types of grammatical errors, using an error correction system based on phrase-based statistical machine translation (SMT) trained on a large scale error-tagged learner corpus. We show that the phrase-based SMT approach is effective in correcting frequent errors that can be identified by local context, and that it is difficult for phrase-based SMT to correct errors that need long range contextual information.

Read the paper · More papers on PaperTik