Evaluating Zero-Shot Large Language Models Recommenders on Popularity Bias and Unfairness: A Comparative Approach to Traditional Algorithms

Gustavo Mendonça Ortega, Rodrigo Ferrari de Souza, Marcelo Garcia Manzato · 2024

Large Language Models (LLMs), such as ChatGPT, have transcended technological boundaries and are now widely used across various domains to enhance productivity. This widespread application highlights their versatility, with a notable presence as recommender systems. Existing literature already showcases their capabilities in this area. In this paper, we present a detailed empirical evaluation of the effectiveness of Zero-Shot LLMs, specifically ChatGPT 3.5 Turbo, without special settings, in calibrating popularity bias and ensuring fairness in movie and TV show recommendations when prompted. We particularly focus on how these models adapt their output, comparing them to traditional post-processing algorithms. Our findings indicate that LLMs, evaluated through metrics such as Mean Average Precision (MAP) and Mean Rank Miscalibration (MRMC), not only perform well but also have the potential to surpass conventional recommender systems models like Singular Value Decomposition (SVD) when paired with calibration methods. The results underscore the advantages of using LLMs in more advanced scenarios due to their ease of implementation and performance.

Read the paper · More papers on PaperTik