Twitter News Classification Using Machine Learning

Muhammad Memoon, Muhammad Umar, Sabha Rani, Muhammad Khalid, Saima Shaheen · 2024

In the era of technology several social media platforms are generating massive data. Social media platforms like Twitter are also generating data in the form of text. Classification of this text is a challenging task because of the diversity in the text, contextual understanding, language complexity and Ambiguity. Classification of news data on Twitter is also an important task in Natural Language Processing (NLP). The news classification task helps us to categorize the Twitter news especially when a topic is trending. In this research, we implement a machine learning algorithm on a new self-created Twitter news dataset. Our methodology involves data collection, preprocessing, feature extraction, and then classification. We evaluate machine learning models including SVM, Random Forest, Naïve Bayes, Logistic regression, and gradient boosting. Experiments show that Random Forest achieved the highest accuracy of 92% on our Los Angeles Twitter News Dataset. This study also provides a new dataset of Twitter news for more advanced text classification approaches.

Read the paper · More papers on PaperTik