Training a CAD classifier with correlated data
Murat Dündar, Balaji Krishnapuram, Matthias Wolf, Sarang Lakare, Luca Bogoni, Jinbo Bi, R. Bharat Rao · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 2007
Most methods for classifier design assume that the training samples are drawn independently and identically from an unknown data generating distribution (i.i.d.), although this assumption is violated in several real life problems. Relaxing this i.i.d. assumption, we develop training algorithms for the more realistic situation where batches or sub-groups of training samples may have internal correlations, although the samples from different batches may be considered to be uncorrelated; we also consider the extension to cases with hierarchical--i.e. higher order--correlation structure between batches of training samples. After describing efficient algorithms that scale well to large datasets, we provide some theoretical analysis to establish their validity. Experimental results from real-life Computer Aided Detection (CAD) problems indicate that relaxing the i.i.d. assumption leads to statistically significant improvements in the accuracy of the learned classifier.