LLM Based Cross Modality Retrieval to Improve Recommendation Performance
Fahad Anwaar, Adil Khan, Muhammad Khalid, Ammad Khalil, Muhammad Awais · 2024
The metadata of items and users play an important role in improving the decision-making process in the Recommender System. In recent times, web scraping-based techniques have been widely utilized to extract explicit user and item meta-data from different social platforms to improve recommendation performance. Currently, Large Language Models (LLMs) have the great potential to replace the traditional web scraping-based paradigm in Recommender Systems. In this paper, we investigated the impact of LLMs and web scraping-based extraction of explicit data on the performance of the Recommender System. Firstly, a cross-modality retrieval-based LLM Gemini is explored to generate semantically enriched textual descriptions of items from digital images. The Gemini LLM is prompted with few-shot prompting on the MovieLens dataset to generate a textual description of movies based on the corresponding movie poster. Secondly, the textual descriptions for each movie in the MovieLens dataset are scraped from the OMDB API. Finally, the cross-modality retrieval-based and scraping-based textual descriptions of items are incorporated into a hybrid Recommender System to assess the quality of explicit data in terms of recommendation performance. The experimental results on the MovieLens dataset demonstrate that LLM-generated content is more effective, achieving a 0.5134 RMSE in enhancing the performance of the Recommender System.