Data Distribution Search to Select Core-Set for Machine Learning

Myunggwon Hwang, Yuna Jeong, Won-Kyung Sung · 2020

This paper contains a strategy to select training data which affects the accuracy of AI positively. When a machine learning (ML) model does not attain a targeted performance, a basic solution for this is to add more data to the model. In this case, we suggest the criteria for selecting more useful data for the learning result instead of adding data randomly. We define a method, data distribution search (DDS), of selecting evenly across all regions based on the distribution of data. In the experiment using MNIST and CIFAR-10, we could confirm that the data set selected by the DDS was superior to a randomly selected set. Ultimately, we could get that there is a data selection method that affects AI performance positively.

Read the paper · More papers on PaperTik