Effect of Domain on Word Prediction quality: Using Stochastic and Neural Network Models
Amiya Kumar Dash, Roshni Pradhan, SARITA TRIPATHY, Santos Kumar Baliarsingh · 2023
Text prediction is the technique of predicting text while the user types. It started with enhancing augmentative and alternative communication and later implemented in different applications like text messages via phones, email systems and many more. This paper showcases the predictive power of models like N-Gram and character buffer when trained on GRU (Gated Recurrent Unit) for different datasets. We have applied two approaches, statistical and neural network approaches which exhibit the neural approach to be the better option when considering feature sets like bigram and trigram. We have investigated the influence domain differences have on word prediction quality and quantified these differences. We have tested this on different domain datasets, each having same corpus size. We found that results show a drop of 26.1% to 41.28% in word prediction accuracy. This work will create more awareness of domain differences in unstructured text data and identify steps to maneuver them.