A Hybrid Approach to Design Automatic Spelling Corrector and Converter for Transliterated Bangla Words
Tanmoy Debnath, Sumaiya Sajnin, Md Montaser Hamid · 2020
In this paper, a spelling corrector and converter has been built for transliterated Bangla words. Transliterated Bangla users are increasing rapidly due to the convenience and general acceptability of this writing style in the internet platforms. For transliterated Bangla, there are no grammatical rules or obligations. Therefore, people generate deviated and ambiguous words. Moreover, people use personal and conventional writing styles in transliterated Bangla. Typographical errors and omission of characters especially vowels for making the words shorter make it more difficult to read transliterated Bangla. To find the solution, we propose a hybrid model that consists of three parts, error detection, error correction, and conversion of transliterated Bangla words. For error detection, we used a dictionary lookup approach and for correction, we used a combination of Damerau-Levenshtein minimum edit distance algorithm and linear search algorithm. Lastly, for conversion, we use the linear search algorithm. For finding out the suitable algorithm for error correction, we tried edit-based, token-based, and sequence-based approaches from text distance library. The model is then evaluated with a dataset comprised of words collected from various internet platforms. The performance of this model shows that it is possible to make an efficient spelling corrector and converter for transliterated Bangla.