Promoting Fairness in Classification of Quality of Medical Evidence
Simon Šuster, Timothy J. Baldwin, Karin M. Verspoor · 2023
Automatically rating the quality of published research is a critical step in medical evidence synthesis.While several methods have been proposed, their algorithmic fairness has been overlooked, even though significant risks may result when such systems are deployed in biomedical contexts.In this work, we study the fairness of two systems with respect to two sensitive attributes: participant sex and medical area.In some cases, we find important inequalities, leading us to apply various debiasing methods.Upon examining the interplay of predictive performance and fairness, as well as medically-critical selective classification capabilities and calibration performance, we find that it is possible to improve fairness through debiasing, but often at a cost to other performance measures.