Assessment of machine learning algorithms for TILs scoring using whole slide images: comparison with pathologists

Arian Arab, Víctor García, Seyed Mostafa Mousavi Kahaki, Nicholas Petrick, Brandon D. Gallas, Weijie Chen · 2024

Machine learning (ML) based whole slide imaging biomarkers have great potential to improve the efficiency and consistency of biomarker quantification, thereby facilitating the development of prognosis models for personalized medicine. Assessment methods in this area are still under-developed. Using the public TiGER (Tumor InfiltratinG lymphocytes in breast cancER) challenge data, we developed a deep neural network-based algorithm for automated tumorinfiltrating lymphocytes (TILs) scoring from whole slide images (WSIs) of biopsies and surgical resections of human epidermal growth factor receptor-2 positive (HER2+) and triple-negative breast cancer (TNBC) patients. The purpose of this study is to assess our model’s performance on a new independent dataset. Seven pathologists independently assessed 320 pre-selected regions of interests (ROIs) across 32 WSIs for TILs scoring. Our results show that there is substantial variability among pathologists in scoring TILs density. We also observed a systematic discrepancy between the ML-based TILs scoring and the pathologists’ manual scoring that led us to develop a calibration between the two. Our calibration reduced the discrepancy, increasing the intra-class-correlation coefficient (ICC) from 0.35 (95% CI [-0.062, 0.625]) for uncalibrated scores to 0.67 (95% CI [0.6, 0.736]) after calibration.

Read the paper · More papers on PaperTik