Comparative analysis of Bangla news classification: a study of fake news detection and multiclass classification using BERT and FastText
Rajeev K. Barua, Mohammed Mahbubur Rahman, Usman Gani Joy · International Journal of Computers and Applications · 2025
In the digital era, the proliferation of misinformation in Bangla news poses a pressing challenge to media credibility. This study provides a detailed examination of Bangla news classification, focusing on two key tasks: binary fake news detection and multiclass news categorization. Leveraging advanced NLP models – BERT and FastText – this research investigates their effectiveness for these tasks and further addresses class imbalance in fake news detection through the application of the Synthetic Minority Oversampling Technique (SMOTE). Without SMOTE, the BERT model achieved a high test accuracy of 0.9879 with an AUC of 0.9676 for binary classification, though its performance on minority classes was limited. With SMOTE, the model's performance improved significantly, achieving an AUC of 0.9997 and a macro F1-score of 0.9834, underscoring the benefits of oversampling for imbalanced data. For multiclass classification, the FastText model attained an accuracy of 89.2% and an F1-score of 0.8919, while BERT recorded an accuracy of 0.8758 and an F1-score of 0.876. This study highlights the value of modern NLP techniques, particularly BERT with SMOTE, in enhancing classification accuracy and combating misinformation in Bangla news.