Performance Differentials in Deployed Biometric Systems Caused by Open-Source Face Detectors

Cynthia M. Cook, Laurie Cuffney, John J. Howard, Yevgeniy B. Sirotin, Jerry L. Tipton, Arun R. Vemury · 2025

Testing of AI systems is important for ensuring accuracy and reliability.In this study, we demonstrate how scenario testing with demographically varied subjects, a form of prospective testing that simulates real-world conditions, revealed significant performance issues in biometric systems prior to broad deployment.Using generalized linear modeling, we show that subjects' measured skin lightness, along with other demographic factors, significantly impacted the probability of failure to detect a face.Failure rates increased from just 0.28% for subjects with the lightest skin in our sample to 24.34% for subjects with the darkest, controlling for other factors.We show that skin lightness, rather than self-reported race, best explained the differences in system performance.We trace these issues to widely used, older methods in open-source packages for face detection.Furthermore, this demographic differential is not observed when testing open-source packages using a different, more curated dataset.Our results highlight the need to evaluate full multi-component, operationally deployed AI systems and the role of scenario testing as a critical component of AI governance.One way to mitigate the likelihood that poor-performing, older open source methods are deployed in an operational system would be to deprecate these functions in favor of higher-performing alternatives.Prospective assessments of AI, in real-world use cases with demographically varied subjects, should be used to identify performance issues before these systems are operationalized.

Read the paper · More papers on PaperTik