Mining Revision Log of Language Learning SNS for Automated Japanese Error Correction of Second Language Learners

Tomoya Mizumoto, Mamoru Komachi, Masaaki Nagata, Yūji Matsumoto · 2011

We present an attempt to extract a large-scale Japanese learners ’ corpus from the revision log of a language learning SNS. This corpus is easy to obtain in large-scale, covers a wide variety of topics and styles, and can be a great source of knowl-edge for both language learners and in-structors. We also demonstrate that the extracted learners ’ corpus of Japanese as a second language can be used as train-ing data for learners ’ error correction us-ing an SMT approach. We evaluate dif-ferent granularities of tokenization to al-leviate the problem of word segmentation errors caused by erroneous input from lan-guage learners. Experimental results show that the character-wise model outperforms the word-wise model. 1

Read the paper · More papers on PaperTik