Comparative Analysis of Bias in LLMs through Indian Lenses
Jhanvee Khola, Shrujal Bansal, Khushi Punia, Rishika Pal, Rahul Sachdeva · 2024
Large Language Models (LLMs) have gained increasing attention in recent times given their exceptional performance in various natural language processing tasks. However, there has been a growing concern that these models may exhibit bias towards certain groups of people. Extensive research has been conducted on LLM bias but most of them operate through a Western standpoint and hence miss details particular to diverse countries in the Global South, like India. This paper aims to investigate the presence of the biases of caste, political and regional sub-culture bias, all in the context of India, in some well-known LLMs. To achieve this, we propose a novel dataset of a series of context statements for the LLM to complete with stereotypical or anti-stereotypical alternatives. The results point out the inherent bias in the generative models which may be due to training over imbalanced datasets or biased human feedback in case of reinforcement learning algorithms used. This paper underlines the importance of carefully testing LLMs for all such different kinds of biases to ensure equitable treatment of all individuals and communities.