Unlocking Deeper Data Insights on Social Media: Removing Hashtag and Tweets Spam for Improved Content Analysis

Gummadi Venkata Nikhil Sai, Robby Aulia Tubagus, Vasala Rohith, Haritha Donavalli · 2024

Users of social media platforms now rely heavily on hashtags to communicate and retrieve messages about certain events, making these platforms an important source of news and information. However, the misuse of hashtags, including spamming and hijacking, can hinder effective communication and lead to the propagation of misinformation. This study proposes a robust methodology for identifying and addressing hashtag spam, emphasizing the crucial significance of data cleansing in ensuring the accuracy and reliability of information. The methodology includes removing punctuation marks and separating duplicate values to create a cleaner and standardized dataset. Hashtags are retrieved using a partially manual spam identifier, which involves human interaction to enhance the overall procedure and improve the precision of the proposed solution. The study also highlights the importance of real-time trend analysis in detecting and removing potentially harmful content, as well as the potential for adaptive filtering mechanisms to dynamically assess the relevance of hashtags within tweets. By integrating both spam and non-spam datasets, the study provides comprehensive insights into social media data and sets the stage for addressing real-time detection, the development of spam tactics, ethical concerns, and monitoring user behaviour. Overall, this research contributes to a nuanced understanding of the challenges of hashtag spam and the importance of data accuracy in social media analysis.

Read the paper · More papers on PaperTik