Statistical Language Models for Spelling Error Detection with Web Search New Word Acquisition
Jui‐Feng Yeh, Guan-Huei Wu, Song-Yi Wang, Chan-Kun Yeh, Yao-Yi Wang · 2021
This paper proposed a statistical language models with internet based new word acquisition for spelling error detection and correction. The statistical language models are composed of off-line and online n-grams. The content in internet is regarded as a large dynamic knowledge resource that is updated by the online users over the world. Due to the word various in internet with rapid renew capability; it is useful for new lexicon learning. The proposed approach can achieve significant improvement in spelling checking. Combining these research manners, the proposed approach is able to achieve the goals of confirming, improving the detection rate of typos. For evaluation, a news dataset with ten categories is gathered as test data. Experimental results show that the proposed approach combining internet web searching outperforms traditional N-gram models, especially for the words that are out-of-vocabulary such as transliteration either in precision rate and recall rate.