A Comparison of Similarity Measures for Neighbourhood Based Collaborative Filtering Recommender Systems
Rohit Beniwal, Rahul Khairwal, Ritvik Mahajan, Sarthak Singh · 2021 Asian Conference on Innovation in Technology (ASIANCON) · 2021
As data continues to grow rapidly at an exponential rate, the importance of data processing systems to extract information is also increasing continuously. Recommender Systems are information processing software that seek to aid decision making by providing relevant suggestions from a given pool of items. In this domain, earlier studies have presented a comparison of similarity measures mostly using Root Mean Squared Error and Mean Absolute Error and to the best of our knowledge, they lacked analysis based on the computation time. Therefore, in this paper, we present a comparison of three similarity measures, namely Pearson Correlation Similarity, Cosine Vector Similarity, and Mean Square Difference on two different datasets known as MovieLens 100k and MovieLens 1M and compared the results using three accuracy metrics namely, Root Mean Square Error, Mean Absolute Error, and Fraction of Concordant Pairs. Our results show that in terms of accuracy, Pearson Correlation Similarity is the best performing similarity measure. Moreover, we also consider the time of simulations for each measure to add another dimension to performance evaluation. As a result, Mean Squared Difference was found to be the fastest among all the three similarity measures