Improvements to Korektor: A Case Study with Native and Non-Native Czech.

Loganathan Ramasamy, Alexandr Rosen, Pavel Straňák · ITAT · 2015

We present recent developments of Korektor, a statistical spell checking system. In addition to lexicon, Korektor uses language models to find real-word errors, detectable only in context. The models and error proba- bilities, learned from error corpora, are also used to sug- gest the most likely corrections. Korektor was originally trained on a small error corpus and used language models extracted from an in-house corpus WebColl. We show two recent improvements: • We built new language models from freely avail- able (shuffled) versions of the Czech National Cor- pus and show that these perform consistently better on texts produced both by native speakers and non- native learners of Czech. • We trained new error models on a manually annotated learner corpus and show that they perform better than the standard error model (in error detection) not only for the learners' texts, but also for our standard eval- uation data of native Czech. For error correction, the standard error model outperformed non-native mod- els in 2 out of 3 test datasets. We discuss reasons for this not-quite-intuitive improve- ment. Based on these findings and on an analysis of errors in both native and learners' Czech, we propose directions for further improvements of Korektor.

Read the paper · More papers on PaperTik