Streamlining Social Music Analytics with Distributed Data Processing: Apache Airflow and MongoDB and Spark SQL

Bhumika Srinivas, Eren Bardak, Ireri Avila, Jessica Brungard, Yihan Cao, Mahesh B Chaudhari · 2025

In the era of data-driven applications, distributed data processing plays a crucial role in managing, analyzing, and deriving insights from vast amounts of user-generated data. This paper presents a scalable data pipeline that uses Apache Airflow, MongoDB Atlas, and Spark SQL to process and analyze Spotify music streaming data. Our pipeline enables a social media-like platform where small friend groups can share and compare their listening history. Using Airflow, we automate data ingestion from the Spotify API, store the data in Google Cloud Storage, and import it into MongoDB Atlas for further processing. We employ Spark SQL on Google Cloud Dataproc to generate analytical insights, such as songs listened to the most, shared music preferences, and user-specific listening patterns. The resulting analytics are then visualized on a dynamic dashboard. This research highlights the effectiveness of distributed data processing frameworks in building scalable and interactive applications while showcasing best practices for integrating cloud-based NoSQL databases with big data analytics tools.

Read the paper · More papers on PaperTik