Next Word Prediction for Urdu using Deep Learning Techniques
Mukhtiar Ahmad Mukhtiar Ahmad, Ali Saeed, M Usman Bhatti, Naveed Hussain, Muhammad Farhat Ullah, Mehmood Anwar · VFAST Transactions on Software Engineering · 2025
A language model for next-word prediction is a probabilistic representation of a natural language that utilizes text corpora to generate word probabilities. These models play a crucial role in text generation, machine translation, and question-answering applications. The focus of this study is to develop an improved algorithm for next-word prediction in Urdu. The study implements deep learning models, including RNN, LSTM, and Bi-LSTM, on a subset of the Ur-Mono Urdu corpus containing 3,000 and 5,000 sentences. To prepare the data for experimentation, tokenization and stemming data cleaning techniques are applied. The study achieved an accuracy of 87% using the RNN model on the first 3,000 sentences of the Ur-Mono dataset and 84% accuracy using the RNN model on the first 5,000 sentences of the Ur-Mono dataset. In conclusion, it can be stated that when the corpus size is small, the RNN outperforms both the LSTM and BiLSTM. However, as the corpus size increases, the Bi-LSTM exhibits superior performance compared to both RNN and LSTM.