Automated Feature Engineering and Hidden Bias: A Framework for Fair Feature Transformation in Machine Learning Pipelines

Rajani Kumari Vaddepalli · ISCSITR-INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ARTIFICIAL INTELLIGENCE AND MACHINE LEARNING · 2025

Automated feature engineering (AutoFE) has become a cornerstone of efficient machine learning (ML), yet its potential to perpetuate or amplify bias remains underexplored.This paper proposes a fairness-aware framework for feature transformation, addressing how AutoFE tools-while optimizing for model performance-may inadvertently encode discriminatory patterns into derived features.Drawing on Ferrario et al. (2022)'s work on bias propagation in ML pipelines and Kamiran & Calders (2019)'s foundational methods for discrimination-aware data mining, we first demonstrate that common AutoFE techniques (e.g., feature synthesis, aggregation) can systematically marginalize underrepresented groups by reinforcing spurious correlations.We then introduce FairFeature, a novel framework that integrates bias metrics (e.g., demographic parity, equalized odds) directly into the feature generation process.Unlike post-hoc fairness adjustments (e.g., adversarial debiasing), FairFeature proactively constrains feature transformations using fairness-aware optimization, ensuring that engineered features meet both predictive utility and equity criteria.Empirical evaluations on real-world datasets (e.g., UCI Adult, COMPAS) reveal that AutoFE without fairness constraints increases disparity by up to 22% in model outcomes, while FairFeature reduces bias by 35-60% with <5% accuracy trade-offs.Our work bridges critical gaps between data engineering and algorithmic fairness, offering practitioners a scalable tool to mitigate 62 hidden biases at the feature level.We further release an open-source library implementing FairFeature to foster adoption.

Read the paper · More papers on PaperTik