Authorship Attribution for Bengali Language Using the Fusion of N-Gram and Naive Bayes Algorithms
DM Anisuzzaman, Abdus Salam · International Journal of Information Technology and Computer Science · 2018
This research shows the authorship attribution for three Bengali writers using both Naï ve Bayes method and a new method proposed by us which performs better than Naï ve Bayes for authorship attribution.Though a lot of works exist in the field of authorship attribution for other languages (especially English); the amount of work in this field for Bengali language is very low.For this experiment, we make our own dataset having 107380 words and 21198 unique words.For both methods, we pre-process our dataset to be compatible to work with the method experiments.For our dataset, Naï ve Bayes gives an accuracy of 86% while our method gives an accuracy of 95%.The main inspiration behind our method is that every author has a nature to write some adjacent words and some single words repeatedly.