Automated Grammatical Error Correction Using Statistical Machine Translation Techniques with Revision Log of Language Learning SNS
Tomoya Mizumoto · Institutional Repositories DataBase (IRDB) · 2015
Recently, natural language processing research has begun to pay attention to second language learning.However, it is not easy to acquire large-scale learners' corpora which are important to a research for second language learner by natural language processing.We present an attempt to extract a large-scale second language learners' corpus from the revision log of a language learning social network service.This corpus is easy to obtain in large-scale, covers a wide variety of topics and styles, and can be a great source of knowledge for both language learners and instructors.I also demonstrate that the extracted learners' corpus of Japanese/English as a second language can be used as training data for learners' error correction using a statistical machine translation approach.For Japanese error correction, we proposed character-based SMT approach to alleviate the problem of erroneous input from language learners.We evaluate different granularities of tokenization to alleviate the problem of word segmentation errors caused by erroneous input from language learners.Experimental results show that the character-based model outperforms the word-based model.For English, I conduct experiments in error correction targeting all types errors using statistical machine translation technique and I analyze the strength and weakness of grammatical error correction using statistical machine translation.I also propose two grammatical error correction methods.One is the method considering multi-word expression.Another is the method using discriminative reranking with POS/syntactic features.I show the effectiveness of multi-word expression and reranking for grammatical error correction.