EuqAud: Detecting Gender Bias in Audio Datasets Using Polynomial Regression-Based Metric

Sriwanthi Jayawardena, Prasanna S. Haddela, Thisara Shyamalee, Amandi Ekanayake, Tharushi Mudalige, Imesha Dhanawardhana · IEEE Access · 2026

With the growing adoption of audio based AI systems in high-stakes domains such as healthcare, law enforcement, and social media, ensuring fairness particularly regarding gender bias has become critically important. While prior work has predominantly addressed disparities in model performance, bias inherent in training datasets remains underexplored. To bridge this gap, we propose EuqAud, a novel, pre-trained and traceable metric that quantifies gender bias in audio datasets using raw acoustic features such as pitch, energy, amplitude, and voice activity. Unlike methods dependent on demographic labels such as race, age or language, EuqAud is designed to be demographic and language agnostic, enhancing its applicability across diverse contexts. The score is computed via polynomial regression with L2 regularization (Ridge regression), yielding robust and generalizable outputs. It spans a range from -10 to 10, where 0 denotes neutral, positive scores indicate male dominant bias, and negative scores reflect female dominant bias. For clarity, bias severity is categorized into three tiers: Neutral (EuqAud6). Evaluation across multiple datasets demonstrates high predictive performance, with R² values between 0.95 and 0.99. By focusing on dataset level bias rather than model outcomes, EuqAud offers a scalable and rigorous solution for advancing fairness in audio-based AI systems.

Read the paper · More papers on PaperTik