Evaluating Bias Detection and Mitigation Approaches Across Classical and Large Language Models

Shriyaa Anand, Aditi Gopinath, Karthikeyan Periyasami · IEEE Access · 2026

In recent years, the increasing reliance on Artificial Intelligence (AI) and Large Language Model (LLM) systems have drawn significant attention to issues of fairness and bias in Machine Learning (ML) models that often inherit disparities present in training data or algorithmic design. IBM’s AI Fairness 360 (AIF360) toolkit provides tools for detecting these biases with fairness metrics and correcting them with mitigation techniques. The mitigation techniques are implemented at three stages: Pre-Processing, In-Processing, and Post-Processing. As no single approach is effective across all scenarios, careful consideration of dataset properties and model behavior is required to mitigate bias for optimizing the model’s performance. This study has taken a closer look at these methods and has compared them in terms of what kind of bias they respond to, amount of bias they can reduce, what degree of accuracy they maintain, and how computationally efficient they are. Through comparing the techniques, the study proposes the best techniques based on the situation, which provides the study with a recommendation on how the techniques can be applied in real-life situations. The research analyses bias reduction in LLMs, including DistilBERT (Bidirectional Encoder Representations from Transformers), GPT-2 (Generative Pre-trained Transformer-2), and BLOOM (BigScience Large Open-science Open-access Multilingual Language Model), assessing fairness measures and mitigation performance on transformer-based models. The applicability of AIF360 is expanded by this integration, which proves the success of the model when it comes to bias reduction among ML models and Natural Language Processing (NLP) systems.

Read the paper · More papers on PaperTik