Supervised and Semi-Supervised Learning for Classifying News Headlines

Md Jabed Hosen, Tanjim Mahmud, Abubokor Hanip, Mohammad Shahadat Hossain · 2025

Due to the linguistic intricacies and dearth of research materials for Bangla, categorizing newspaper headlines in this language poses particular hurdles. This study investigates the effectiveness of supervised and semi-supervised learning strategies for categorizing Bangla news headlines into pre-established groups. Support Vector Machine (SVM) and Random Forest, two supervised learning algorithms, and FixMatch and Generative Adversarial Networks (GAN), two semi-supervised techniques, were employed. According to our experimental results, the Random Forest model achieved an accuracy of 66%, while the SVM model reached 69%. The FixMatch model, utilizing both labeled and unlabeled data, demonstrated an accuracy of 68%, slightly outperforming traditional approaches. The GAN-based semi-supervised model also showed promising results, with an accuracy of 67%. Our findings suggest that semi-supervised learning methods can improve model performance, even when all the data is fully labeled, particularly when additional techniques such as data augmentation or other supplementary information are utilized. This study provides valuable insights into various learning paradigms for classifying Bangla text, contributing to the advancement of natural language processing for low-resource languages.

Read the paper · More papers on PaperTik