Modular post-hoc enhancements for zero-shot histopathology classification using vision-language models

Shahd M. Noman, Mayar Tarek Henedak, Mustafa Elattar · Machine Learning Science and Technology · 2025

Abstract Zero-shot learning with vision-language models (VLMs) has shown promising results in histopathological image classification. However, existing approaches often under-utilize domain-specific adaptations that could enhance performance in specialized medical tasks. In this study, we systematically evaluate three modular, post-hoc enhancements to standard zero-shot VLM pipelines: (1) cosine similarity calibration for affinity matrix construction, (2) prompt template enhancement via diverse clinical prompts, and (3) adaptive hyperparameter tuning using a Gaussian mixture model-based method for adjusting pseudo-labeling weights (λ) and support set contributions (γ), following Zanella et al We conduct extensive zero-shot experiments on the NCT-CRC-HE-100 K dataset, using 60 000 patches for validation of inference-time hyperparameters and reserving 40 000 patches as an independent test set. No training or fine-tuning was performed. The evaluation was carried out across five VLMs: CLIP, CONCH, PLIP, Quilt-B16, and Quilt-B32. Results show that each enhancement module yields significant accuracy gains. Cosine calibration improves CLIP (+7.8%), CONCH (+1.53%), PLIP (+14.3%), Quilt-B16 (+16.1%), and Quilt-B32 (+11.4%). Prompt enhancement yields accuracy gains in PLIP (+16.8%) and Quilt-B32 (+29.9%), with limited effect on CLIP or CONCH. Adaptive tuning yields large improvements in PLIP (+28.4%). The best performance is achieved with the full combination of enhancements, with PLIP reaching 87.88% (+27.88%, p < 0.001) and Quilt-B32 reaching 88.38% (+36.1%) over their baselines. CLIP showed significance only with cosine calibration (p < 0.001), reflecting its contrastive learning backbone. Although CONCH reached the highest raw accuracy (92.63%) under the full enhancement setting, this performance appears primarily driven by its histopathology-specific pretraining; statistically significant gains emerged only in cosine-based combinations, indicating limited responsiveness to prompt or tuning. Our findings highlight the critical role of structured modular enhancements in optimizing VLM performance for specialized clinical domains.

Read the paper · More papers on PaperTik