Detection and Correction of Real-Word Errors in Bangla Language

Md. Mashod Rana, Mohammad Tipu Sultan, M. F. Mridha, Md. Eyaseen Arafat Khan, Md. Masud Ahmed, Abdul Hamid · 2018

Detection of spelling error is not so facile in Bangla. To check for real-world error in a sentence, it comes with more difficulties. In this paper, we focus on correcting homophone error in real-word error. We use N-gram Model which is used in many purposes like machine translation, speech recognition, to extract syntactic information etc. We have used a combination of Bi-gram and Tri-gram with candidate word which is going to be detected whether it is a real-word error or not. We have developed corpora which contain: (i) one of them is a collection of sets of homophone (confusing) word, (ii) another two are the collection of bigrams and trigrams using homophone word and (iii) other seven are the test sets. A candidate word extracts the set of homophone words from the corpus. In our proposed method, we create tri-gram and bigram using homophone word, then it checks the validity and takes the frequency of bi-gram or tri-gram, and finally calculates the probability for making the final decision about the candidate word. We have used around a million words to inspect our system. Our proposed method achieves more than 96% accuracy in detecting and correcting real-word errors of Bangla Text.

Read the paper · More papers on PaperTik