FairPreprocessor: Better Fairness Via Addressing Imbalanced Data Through Synthetic Data Generation and Mitigating Biased Labels

Hem Chandra Joshi, Sandeep Kumar · IEEE Intelligent Systems · 2025

The machine learning(ML) model acquires logic from the training dataset, and any bias within it impacts the model’s decision. Previous studies have revealed that ‘biased labels’ and ‘imbalanced data’ are significant causes of bias in the training dataset. This study proposes a preprocessing approach, FairPreprocessor, that addresses ’imbalanced data’ through rebalancing the internal data distribution by employing synthetic data techniques grounded in differential evolution. It also selects the most suitable crossover rate in synthetic data generation to achieve better fairness. Additionally, it identifies and removes biased labels through situation testing, thereby mitigating their effects and developing fairer ML software. To facilitate an open science, this study’s source code and datasets are available online athttps://github.com/sendgmale/FairPreprocessor.

Read the paper · More papers on PaperTik