Prediction of Semantically Correct Bangla Words Using Stupid Backoff and Word-Embedding Model

Tanni Mittra, Linta Islam, Deepak Chandra Roy · 2019

Word prediction is an essential technique used in different text entry environment to facilitate error-free writing. It also used as a helping hand for people with different types of disabilities. Word prediction technique is available in different languages. But developing an optimized Bangla word predictor is still a great research challenge. To overcome the challenge we propose a hybrid method to predict Bangla words. The stupid backoff language model is used to detect the most-probable words that may fit into the previously typed sentence by calculating the word sequence frequency. The novelty of this work is that it can provide the semantically correct words as a suggestion. The Word-Embedding model is used to maintain the semantic context of the word. To test this approach, a large corpus is built consisting of almost 0.5 million data. We compared our approach with other well-established methods. The proposed methodology surpasses them by obtaining 83% accuracy. The approach is also computationally efficient as the running time is linear with the prediction length.

Read the paper · More papers on PaperTik