Data Clustering for Training Set Selection
Christopher Bowman, Duane DeSieno, Christopher R. Tschan · Infotech@Aerospace 2011 · 2011
This paper addresses issues related to building usable size training sets for training of neural networks where there is a massive amount of data. Neural networks typically have the characteristic that more data is better for training and generalization. However, for some problems, getting data is not a problem and in fact becomes an issue when there are massive amounts of data. This is particularly important when the training set contains infrequently occurring examples that need to be learned. When training sets get large enough that the time to do the training is excessive, one needs to select a subset of the data for training. The issue then becomes “What subset of data should be used to best learn all the desired behaviors?”