Machine Learning for Text Classification on Twitter: A Literature Review
Muneer Hazaa Alsurori, Ahlam Enan, Rahuf Alwan, Wafa Algumaei, Somia Alturki, Entsar Alkahtany · American Journal of Data Mining and Knowledge Discovery · 2023
This literature review examines the application of machine learning (ML) techniques for text classification on Twitter. With the immense volume of data generated on social media platforms like Twitter, there is a need for automated methods to extract valuable information. ML, known for its ability to learn patterns and relationships in large datasets, has gained significant attention in this context. The purpose of this review is to explore the background and aim of ML for text classification on Twitter, the methods employed, the results obtained, and the conclusions drawn. The review begins by discussing the background and aim, emphasizing the vast amount of data available on Twitter and the need for automated techniques to extract useful information from this data. It highlights the significance of ML in addressing this challenge, particularly in tasks such as sentiment analysis, topic modeling, and spam detection, which play a crucial role in social media analysis. Next, the review provides an overview of the methods used in various studies on text classification using Twitter data. It explores the latest approaches and techniques employed in ML, including feature extraction methods like bag-of-words, n-grams, and word embeddings. It also discusses the preprocessing steps involved in preparing Twitter data for classification tasks. subsequently, the review presents the results obtained from different studies in the field. It discusses the performance metrics used to evaluate the effectiveness of ML models, highlighting measures such as accuracy, precision, recall, and F1-score. The review also discusses variations in performance across different classification tasks, providing insights into the strengths and limitations of the approaches used.