An Approach for Validating Quality of Datasets for Machine Learning

Junhua Ding, Xinchuan Li · 2018

There are basically two ways for improving the accuracy of machine learning: building relevant machine learning models, and providing high quality datasets for training the models. Significant efforts have been made in designing powerful machine learning models. Furthermore, many open-source datasets have been created for machine learning research. However, research on assessing the impact of the quality of a dataset on the accuracy of a machine learning system has not received attention. In this paper, we present an experimental study to show how the quality of datasets impact the accuracy of machine learning models. We discovered a common problem in datasets that could greatly impact the accuracy of machine learning. This problem could also exist in many other machine learning systems, especially those that are developed using crowd-sourced datasets. This problem is difficult to detect using traditional validation approaches. We propose a novel technique based on metamorphic testing for validating a machine learning system together with its training and testing data. The key to metamorphic testing is to create tests that will adequately test the system. We propose an approach for creating such tests. The effectiveness of the proposed approach is demonstrated through a case study of automated classification of biological cell images.

Read the paper · More papers on PaperTik