A Combinatorial Approach to Reduce Machine Learning Dataset Size

Megan Olsen, Mohammad S. Raunak, D. Richard Kuhn, Hans Van Lierop, Fenrir Badorf, Francis Durso · 2025

Although large datasets may seem to be the best choice for machine learning, a smaller dataset that better represents the important parts of the problem will be faster to train and potentially more effective. Although dimensionality reduction focuses on reducing features (columns) of data, dataset reduction focuses on removing data points (rows). Typical approaches for reducing dataset size include random sampling, or using machine learning to understand the data. We propose using combinatorial coverage and frequency difference (CFD) techniques from software testing to choose the most effective rows of data to generate a smaller training dataset. We explore the effectiveness of four approaches to reduce a dataset using CFD, and a case study showing that we can produce a significantly smaller dataset that is more effective in training a Support Vector Machine than the original dataset or datasets generated by other approaches.

Read the paper · More papers on PaperTik