Next Word Prediction in Bangla Using Hybrid Approach

S. M. Nuruzzaman Nobel, Shirin Sultana, Md All Moon Tasir, Md. Saifur Rahman · 2023

The impact of language models in various applications like machine translation, speech recognition, and chatbots has transformed text-based services. Despite being spoken by over 300 million people, the Bangla language has not yet developed advanced language models due to its unique linguistic traits and limited resources. Traditional language models struggle with the intricacies of Bangla’s linguistic traits, hindering the development of advanced prediction systems. This study aims to bridge this gap by introducing an approach focusing on autocompletion and sequence prediction. The study proposes the integration of the Trie data structure, Convolutional Neural Network (CNN) with Long Short-Term Memory (LSTM), and N-gram methodologies for Bangla word completion and sequence prediction. The Trie data structure stores the Bangla vocabulary and extracts words based on user-entered prefixes, improving efficiency and preventing misspellings. The hybrid architecture is designed to capture long-range relationships and comprehend contextually rich patterns. This makes it particularly efficient at handling complex compound words, rich inflections, and variable word order. The experimental evaluation involves a comprehensive dataset comprising approximately 50 thousand Bangla language samples collected from diverse sources, ensuring a representative linguistic spectrum. The results show the potential of this approach to revolutionize Bangla search engines, keyboards, and recommendation algorithms.

Read the paper · More papers on PaperTik