Large-scale recommender systems using Hadoop and collaborative filtering: a comparative study
Mohamed El Amine Chafiki, Oumayma Banouar, Mohamed Benslimane · Mathematical Modeling and Computing · 2024
With the rapid advancements in internet technologies over the past two decades, the amount of information available online has exponentially increased. This data explosion has led to the development of recommender systems, designed to understand individual preferences and provide personalized recommendations for desirable new content. These systems act as helpful guides, assisting users in discovering relevant and appealing information tailored to their specific tastes and interests. This study's primary objective is to assess and contrast the latest methods utilized in recommender systems within a distributed system architecture that relies on Hadoop. Our analysis will focus on collaborative filtering and will be conducted using a large dataset. We have implemented the algorithms using Python and PySpark, enabling the processing of large datasets using Apache Hadoop and Spark. The studied approaches have been implemented on the MovieLens dataset and compared using the following evaluation metrics: RMSE, precision, recall, and F1 score. Their training times have also been compared.