Training with Probability Map: an Effective Framework for Deep Learning Training with Multi-observer Labeled Datasets
Hossein Arabi, Habib Zaidi · 2021 IEEE Nuclear Science Symposium and Medical Imaging Conference (NSS/MIC) · 2021
The substantial inter-observer variability makes grading prostate tissue micro-array (TMA) cores to determine Gleason scores difficult. Due to the huge variance in pathologists' scores and the lack of proven ground truth, developing a deep learning-based model for automated grading of prostate TMAs is extremely difficult. In light of this, a number of ways for training deep learning algorithms with noisy or multi-observer labeled datasets have been developed. This paper evaluates a number of widely used and effective methods for training a model with a heterogeneous dataset. In addition, deep learning model training with a probability map is presented as a solution to this problem. To test the alternative training procedures, 240 prostate TMA cores from the Gleason2019 challenge (evaluated independently by six pathologists with Cohen's Kappa agreement coefficients between 0.40 and 0.73) were used. Training with single pathology scores, training with a single label map generated by majority voting (MV) and STAPLE methods, training the model with L1-norm and L2-norm loss functions using the entire scores of pathologists, and training with a single probability map obtained from the entire scores of pathologists are among the strategies evaluated in this study. The ground truth was the label map created by STAPLE. The model trained with the proposed probability map had the best accuracy of 0.93, whereas the models trained with the L2-norm and STAPLE had accuracy of 0.89 and 0.86, respectively. In comparison to these models, other approaches produced poor results. In comparison to the commonly utilized methodologies, the proposed simple training strategy yielded the highest accuracy. A multi-observe labeled dataset could benefit from the proposed approach for decision fusion and ground truth estimation.