How Data Diversification Benefits the Reliability of Three-version Image Classification Systems
Mitsuho Takahashi, Fumio Machida, Qiang Wen · 2022
Recently, we have witnessed an increased use of systems employing Machine learning (ML) models. The dependability of such systems is significantly impacted by the inference results of ML models that may not always be correct. Nversion ML systems can be adopted for improving output reliability by detecting or correcting errors by combining multiple inference results. In this study, we focus on a three-version ML system that combines the inference results of an ML model on diversified input data to make the system reliable for ML image classification tasks. The three-version image classification system uses only one deep-learning classifier while generating three inference results in response to diversified inputs. We use several image transformation methods to generate diversified input data for inferences. The system reliability is evaluated by the coverage of errors and the certainty of accurate predictions, as those metrics are decision method agnostic. We confirm that appropriately combined inference results from diversified data can increase the coverage of errors and improve reliability while maintaining the certainty of accurate predictions. In order to search for such effective combinations of diversification methods, we propose the neuron coverage improvement rate (NCIR) as an indicator of data diversity. Through the experiments, we show that the NCIR tends to have correlations with the coverage of errors and the certainty of accurate predictions, indicating the usefulness of the indicator.