Sentiment Analysis of Translated BBC Hindi News Articles Using Machine Learning: A Comparative Study of Classification Algorithms
Vishnu Achutha Menon, T. K. Sateesh Kumar, Juby Thomas, Lijo P Thomas · 2024
News consumption in India has shifted significantly with the rise of Web 2.0 technologies, leading to the prominence of platforms like BBC Hindi News, which blends international credibility with regional relevance. The dataset comprises over 5000 articles from BBC Hindi (Kaggle), with sentiment distributed as 97% neutral, 2.1% positive, and 0.7%negative. Text preprocessing steps include cleaning, tokenization, stopword removal, and lemmatization. Features were extracted using TF-IDF vectorization, and two machine learning models-Logistic Regression and Random Forest-were employed for classification. Performance metrics such as accuracy, precision, recall, F1-score, and confusion matrices were used for evaluation. Logistic Regression achieved an accuracy of 97.38%, whereas Random Forest marginally outperformed with 97.99%. Random Forest also demonstrated better handling of minority classes, reflected in higher macro-average metrics compared to Logistic Regression. Despite the improvements, both models exhibited a strong bias toward the dominant neutral class, with difficulties in accurately classifying positive and negative sentiments. The study highlights the importance of addressing class imbalance through techniques like oversampling minority classes or adjusting class weights. Random Forest's performance suggests ensemble techniques are better suited for imbalanced datasets, though further optimization through hyperparameter tuning and feature refinement is recommended.