A statistical approach to quality evaluation of AI mislabel detection algorithm

Jiayi Lian, Kexin Xie, Kevin Choi, Xueying Liu, Balaji Veeramani, Sathvik Murli, Alison Hu, Laura Freeman, Edward Bowen, Xinwei Deng · Quality Engineering · 2025

Testing and evaluating the quality of Artificial Intelligence (AI) algorithms is important for confidently deploying them in real applications. AI algorithms depend both on the hyper-parameters and the nature of the training data. A measure of the quality of an AI algorithm can be obtained by exhaustively evaluating the performance of the algorithm across hyper-parameters and data quality factors. However, such a procedure is challenging as it is not practical to enumerate all possible level combinations of the hyper-parameters and data quality factors. In this work, we present a principled framework using a statistical approach to systematically conduct quality evaluation of AI algorithms, named as QE-AI. The proposed framework consists of an efficient space filling design in a high-dimensional constraint space and an effective surrogate model using an additive Gaussian process to enable efficient quality evaluation of AI algorithms. We demonstrate the performance of the proposed approach for an AI mislabel detection algorithm.

Read the paper · More papers on PaperTik