Pipilika N-Gram Viewer: An Efficient Large Scale N-Gram Model for Bengali

Adnan Ahmad, Mahbubur Rub Talha, Md. Ruhul Amin, Farida Chowdhury · 2018

In this paper, we introduce a large-scale Bengali N-gram model, trained on online newspaper corpus and present results and analysis of two different experiments done by using the model, namely Context-aware spell checker and Trending topic detection. We also present the process with emphasis on the problems that arise in working with data at this scale. One signicant aspect of our N-gram model is that the model contains information of N-gram occurrence per day over a period of eight years, from the year 2009-2017. This enables further applications of the model, for example, Trending topic detection. Our Bengali N-gram language model contains N-grams up to 5-gram with more than 2 million unique Unigrams and over 656 million duplicate Unigrams. We evaluate our model by calculating the perplexities of different years. We obtain F -score of 86.6% in an experiment of Context-aware spell checker. In another experiment, we successfully detected the trending topics of a given time frame. This paper also presents first Bengali N-gram viewer, where one can query by a particular N-gram and see the resulting graph of the frequency of that term occurred using different time frames.

Read the paper · More papers on PaperTik