Learning from Multiple Annotators: Hierarchical Deep Learning Training Scheme for Prostate Gleason Cancer Grading

Hossein Arabi, Habib Zaidi · 2021 IEEE Nuclear Science Symposium and Medical Imaging Conference (NSS/MIC) · 2021

The grading of prostate tissue micro-array (TMA) cores is prone to high inter-observer variability. When it comes to determining Gleason scores, there is a lot of disagreement among pathologists. In this context, merging extremely varied decisions made by multiple observers is a problem for developing a machine learning model for automated scoring of prostate TMAs. In the absence of a known ground truth, a machine learning strategy that can reliably mod-el/discover the underlying agreement/true labels from various observers is critical. In this context, this paper proposes a hierarchical training approach for developing a deep learning model that can effectively decode/discover real labels from a multi-observer labeled dataset. Moreover, the proposed method (DL-Hrch) is compared with a number of commonly used approaches for dealing with noisy or multi-observer labeled datasets, such as training based on the confusion matrix (DL-Cmat). The DL-Hrch framework entails training of a deep learning model successively starting from least-agreed to the most-agreed data points. In this scheme, the model would be fine tunned with the consensus of the entire annotators at the latest iterations. For evaluation, 240 prostate TMA cores, graded independently by six observers (from Gleason2019 challenge), were employed. Cohen’s Kappa agreement coefficients among the different observers were between 0.40 and 0.73. The label maps generated from the STAPLE algorithm were regarded as ground truth. The DL-Hrch model exhibited the superior accuracy of 0.94 in terms of Dice index, while the DL-Cmat resulted in the Dice index of 0.92. The other decision fusion algorithms exhibited greatly inferior accuracy compared to the DL-Hrch approach. The proposed hierarchical training scheme yielded the highest accuracy in comparison with the commonly used algorithms to deal with noisy or multi-observer labeled datasets. The proposed training scheme could be utilized for decision fusion from multiple observers and estimation of the true label maps.

Read the paper · More papers on PaperTik