Distribution-Based Splitting of Datasets
Timm J. Peter, Dennis Kekec, Oliver Nelles · 2018
The problem of data splitting is discussed and two novel approaches to split datasets based on an efficient evaluation of estimated probability density functions (pdf) are introduced. Due to their structure, the approaches are capable of selecting subsets, that approximate the input point distribution of the original dataset. The first method fills the datasets alternatingly with data points. The second method allocates the points to the dataset with a probability proportional to their fit. Both approaches base on rough a simplistic way of evaluating pdfs, they are estimated using approximate kernel density estimation. Additionally, an subsequent optimization of the point distribution is introduced, which also relies on pdf estimation. Since datasets are split based on the point distribution of the original dataset, the estimation of pdfs presents a more sophisticated criterion to base a data splitting concept on, compared to existing procedures. In order to validate the functioning of the approaches, their performance is analyzed for two artificial datasets, comparing it to an existing method and the mean performance by random splits. Both approaches and the subsequent optimization lead, on average, to an improved performance. The first method shows the most promising performance improvement.