Movie Recommendation System Using Euclidean Distance
Muneeb Sami Khan, Muhammad Zunnurain Hussain, Muhammad Irtaza Amaad, Muhammad Zulkifl Hasan, Aqsa Khalid, Saad Hussain Chuhan, Rimsha Awan, Muhammad Atif Yaqub, Farhan Ashraf · 2024
In this study, we pivot from the commonly used cosine similarity to the Euclidean distance as the foundation for our film recommendation algorithm. Initially, data pertaining to users, films, and ratings are loaded using the panda’s library. The algorithm then computes normalized ratings, tabulates the number of films in each genre, and visualizes this distribution. The frequency of ratings for each genre is similarly plotted. A dictionary is maintained to store movie IDs for each genre. These values are then integrated into the normalized rating matrix. For each genre, the normalized ratings of movies are cataloged in a lexicon. Using the normalized ratings matrix and movie IDs, a Data Frame is generated. Subsequently, the movie and ratings datasets are merged. A pivot table captures the number of films each user has watched per genre. The study also be happy with their purchase if these recommendations are tailored to their needs. As a result, the client would utilize this application once more. Because of the significant money gained determines the average number of films a user watches in each genre. This pivot table is refined to display movies in each genre that a user has watched, but only those exceeding the genre’s average viewership. Instead of cosine similarity, the Euclidean distance is computed to gauge the similarity between movies, offering a fresh perspective on recommendation systems.