Conformal Prediction for Deep Learning Classification Model in Histopathological Images
Santiago Gelvez, Diego Clavijo, David Romo‐Bucheli · 2025
Deep learning has shown significant results in histopathology image analysis. However, despite their high pre-dictive accuracy, conventional machine learning (ML) models often lack mechanisms to quantify uncertainty, which is necessary for building trust in clinical scenarios. This limitation hinders the integration of these models into real-world diagnostic workflows. Conformal Prediction is a relatively recent statistical frame-work that enhances ML models by generating prediction sets according to a defined confidence level. These sets indicate the range of plausible outcomes for each input, thus allowing clinicians to assess model confidence. In this study, we explore Conformal Prediction methods for multi-class classification of histopathological images using state-of-the-art convolutional neural networks, specifically DenseNet and EfficientNet. We implement three Conformal Prediction approaches: Inductive Conformal Prediction(ICp)with and without Adaptive Prediction Sets (APS), and Mondrian Conformal Prediction (MCP). Each method is evaluated on a prostate cancer whole-slide images from two different institutions. Evaluation metrics include the obtained coverage, N-criterion, Fairness in Subgroup Coverage (FSC), and Stratified Size Coverage(SSC), which assess adaptability across institutional and predictive set variations. Our results show that all the methods were able to generate predictions sets with coverage close to the desired confidence levels$(\alpha=0.9)$. ICP and MCP obtained similar coverage levels (around 0.88-0.89) while maintaining more compact prediction sets when compared to APS. Additionally the FSC and SSC metrics show that the conditional coverage was approximately equal, with the lowest result corresponding to the DenseNet architecture 0.86. Contrary to our expectations, MCP did not outperform ICP, suggesting that additional information associated to clinical site did not translate into clear improvements. Regarding the convolutional architectures, DenseNet was able in average to achieve a better N-criterion than the EfficientNet. This is notorious for both ICP (1.76 vs 2.01) and MCP (1.78 vs 2.01) methods. Our results show that the APS conformal method outperforms ICP and MCP for this classification task. Additionally, the selection of the architecture might introduce important differences in the uncertainty of the predictions as reflected in the N-criterion values of the resulting conformal prediction sets.