Training Set Class Distribution Analysis for Deep Learning Model - Application to Cancer Detection

Ismat Ara Reshma, Margot Gaspard, Camille Franchet, Pierre Brousset, Emmanuel Faure, Sonia Mejbri, Josiane Mothe · Open Archive Toulouse Archive Ouverte (University of Toulouse) · 2019

Deep learning models specifically CNNs have been used successfully in many tasks including medical image classification. CNN effectiveness depends on the availability of large training data set to train which is generally costly to obtain for new applications or new cases. However, there is a little concrete recommendation about training set creation. In this research, we analyze the impact of different class distributions in the training data to a CNN model. We consider the case of cancer detection task from histopathological images for cancer diagnosis and derive some useful hypotheses about the distribution of classes in the training data. We found that using all the training data leads to the best recall-precision trade-off, while training with a reduced number of examples from some classes, it is possible to inflect the model toward a desired accuracy on a given class.

Read the paper · More papers on PaperTik