Real-time Tweets Analysis using Machine Learning and Bigdata

P Nandieswar Reddy, Sai S, Rithvika Alapati, D. Radha · 2024

Twitter's exponential increase in data has transformed it into a rich source for machine learning research, unveiling patterns of opinions and behaviors. This research introduces an automated pipeline built on Kafka, adept at handling large volumes of real-time Twitter data. Utilizing the Twitter API, tweets bearing the hashtag #justice are retrieved and efficiently processed via Kafka. PySpark facilitates data analysis by enabling aggregation, Jaccard similarity computation, and k-means clustering. Preprocessing with Python libraries and TF-IDF enhances analysis accuracy. Jaccard similarity aids in comparing tweets across socio-economic categories. Aggregated data is stored in MongoDB for further examination. Additionally, Streamlit offers visualization tools for effortless exploration of Twitter discourse trends. Thus, this project exemplifies the potential of big data and its impact on investigating social justice discourse in the digital age.

Read the paper · More papers on PaperTik