Natural Language Processing based Stochastic Model for the Correctness of Assamese Sentences

Manash Pratim Bhuyan, Shikhar Kumar Sarma, Mirzanur Rahman · 2020

Increasing the number of social media users and the way these users write the regional languages have pushed the researchers to think about the originality of these languages. The native speakers are in fear that with time the regional languages may lose their originality. For such reason, the correctness checking at the time of writing is an important task in the field of natural language processing. The conventional methods for checking the correctness of a sentence are normally carried out by applying various syntactic and semantic rules and the rules are endless because of the free word order nature of the Assamese language. On the other hand, in the computational filed, the data-driven models are more appropriate than the rule-based system and the data-driven model is free from the rules. In this research work, a data-driven model is designed by using the different types of n-gram models like unigram, bigram, and trigram to check the correctness of the Assamese sentences along with the linear interpolation method. Assamese is one of the twenty-two official languages of India, predominantly spoken in Assam and its neighboring states. A corpus of the size of around 1 million words is used to train the system. During testing, four different levels of experiments are carried out one for the correct sentences and incorrect sentences. The F1-score of the system in correcting the users' sentences is above 60%.

Read the paper · More papers on PaperTik