Myanmar Spelling Checking and Correction Based On Noisy Channel Model
Saung Pwint Phyu, Win Lelt Lelt Phyu · 2024
Spelling correction is a crucial task in natural language processing, enabling improved understanding and processing of text by correcting phonetic errors, context errors and typographical errors. This paper uses Noisy Channel Model that provides the probabilistic framework for spelling correction by treating the process of typing as a transmission over a noisy channel, where the intended word is distorted. According to the Bayesian inferences, this model consists of two main components: the language model, which captures the prior probability of word sequences, and the error model, which estimates the probability of spelling errors given the intended word. The data are extracted from UCSY Monolingual Corpus. The Spell-error Corpus is manually created for training the error model. The language model is trained using the Spelling Correction Corpus, which is manually segmented and corrected from the UCSY Monolingual Corpus. Damerau-Levenshtein Distance Algorithm is used for candidates generation. Experimental findings indicate that the spelling correction method effectively addresses each type of error, achieving an overall accuracy exceeding 86% across all error categories.