A Research on the Japanese Open-source Automatic POS Taggers
Mao Wen-wei · Computer-assisted Foreign Language Education · 2012
The automatic POS tagging technology has matured to provide a strong support for the corpus building.Unlike the native speaker's corpus,the learner's outputs are flooded with errors.This will definitely interfere with the accuracy of the tagging.Therefore,in addition to accuracy,the anti-interference ability should also be taken into account.This paper focuses on the Japanese open-source automatic POS taggers,calculates the accuracy when they are used to tag a group of the learner's texts and observes whether the performance are affected by the quality of texts.Results of the study indicate that MeCab is the best and ChaSen acts better than JUMAN.It is also proved that the accuracy of the learner's corpus tagging is even better than the performance when they are used to tag the native speaker's corpus.Therefore,the taggers can be used as a powerful tool during the construction of learner's corpus.