Twitter Data Analysis Using Apache Streaming

Lavanya K. Sendhilvel, Kush Diwakar Desai, Simran Adake, Rachit Bisaria, Hemang Ghanshyambhai Vekariya · Advances in computational intelligence and robotics book series · 2022

Real-time data from social network sites like Twitter or Facebook has been a popular source for analytics and researchers in the recent years due to various factors like large amount of data, structured-ness, and popularity. Analyzing data is a very common requirement today, but such requirements become difficult when there is a bulk of data which needs to processed and analyzed in real time. Analyzing large number of tweets from Twitter to get different patterns and extract useful information is a massive challenge. Apache Spark is a platform that can be used to handle big data efficiently, and it offers faster solutions compared to Hadoop. This chapter addresses the issue of real-time analyzing and filtering the tweets as per the user's requirements from among the millions of other streaming tweets and classifies them into various categories. It creates an interactive automatic system that splits data based on important keywords and displays a graphical representation of connected tweets using Apache Spark.

Read the paper · More papers on PaperTik