Re: Deep learning outperformed 11 pathologists in the classification of histopathological melanoma images

Cyrill Géraud, Klaus Griewank · European Journal of Cancer · 2019

We read with great interest the recent publication by Hekler et al. [[1]Hekler A. Utikal J.S. Enk A.H. Solass W. Schmitt M. Klode J. et al.Deep learning outperformed 11 pathologists in the classification of histopathological melanoma images.Eur J Cancer. 2019; 118: 91-96Abstract Full Text Full Text PDF PubMed Scopus (124) Google Scholar]. The authors compared analysis of histology images from melanomas and nevi by deep learning with assessment by 11 humans. The results are interesting; however, we have considerable concerns regarding the title, study design and the conclusions made. Reviewing the invitations we received for participation (provided to the editors), this was sent to a considerable number of physicians including resident physicians without board certification in dermatology, dermatopathology or pathology by email. The list of people invited and the listed study participants we know of already exceeds the 14 reported in the manuscript. As only eight of the 11 study participants were board-certified, and the manuscript reports two participants performed on par with deep learning, the study title ‘outperformed 11 pathologists’ is misleading. Participant information was based on a freely accessible google form not restricted to invited participants making verification impossible. The 20 min allotted in the invitation are insufficient for a pathologist to assess 100 images of melanocytic lesions with the required prudence (=12 seconds per image). The senior author's invitation letter explicitly stated that the study design, applying cropped images, provided image material inappropriate for a human pathologist, the hypothesis being deep learning could better distinguish nevus and melanoma under these circumstances. We find it problematic that the manuscript makes statements such as ‘Deep learning outperformed 11 pathologists’ and ‘The aim of this study is to perform such a first direct comparison’ without explicitly stating the caveat mentioned in the invitation letter that the study design intentionally impeded diagnosis by a human pathologist. It is also scientifically questionable to conceive studies with the primary aim of showing superiority of a new technique. A major issue of the Hekler et al. study is applying cropped images showing only a small random portion of the actual lesion. No dermatopathologist would make a real routine patient diagnosis being able to inspect only a small piece of one section of the entire lesion. Melanomas can be highly heterogeneous or nevus-associated which is why their diagnosis (at least for humans), relies on a number of criteria that can only be assessed when the complete lesion is visible. A key choice a pathologist has when facing incomplete melanocytic lesions is deciding whether the visible features are insufficient for diagnosis, and further work-up or re-excision of the complete lesion is required. Unfortunately, the study design did not provide this option. We personally opted against participation in the study feeling that for most images included the binary approach forcing participants to choose a diagnosis of nevus or melanoma was problematic. One can argue deep learning algorithms may recognise specific signs not readily recognised by a human pathologist. However, reviewing the 100 histology pictures included in the trial, we identified more than 15 (>15%) having no recognisable melanocytic lesion whatsoever presumably representing perilesional normal skin tissue. Certainly, a human pathologist cannot make a distinction of nevus or melanoma if a melanocytic lesion is not present in the image demonstrated. If Hekler et al. believe deep learning can distinguish nevus from melanoma based solely on perilesional tissue, they need to show specific data to support this theory. The authors should provide access to all cropped images with the corresponding diagnoses and distribution of choices made by deep learning and participants to allow the readership an unbiased assessment of the data. Interestingly, despite citing discordance rates among expert histopathologist of 25–26% [[1]Hekler A. Utikal J.S. Enk A.H. Solass W. Schmitt M. Klode J. et al.Deep learning outperformed 11 pathologists in the classification of histopathological melanoma images.Eur J Cancer. 2019; 118: 91-96Abstract Full Text Full Text PDF PubMed Scopus (124) Google Scholar,[2]Hekler A. Utikal J.S. Enk A.H. Berking C. Klode J. Schadendorf D. et al.Pathologist-level classification of histopathological melanoma images with deep neural networks.Eur J Cancer. 2019; 115: 79-83Abstract Full Text Full Text PDF PubMed Scopus (106) Google Scholar], Hekler et al. performed studies where melanoma or nevus were defined by a single histopathologist. The discordance rates cited are problematic as they predominantly refer to difficult to classify melanocytic lesions, not included in the Hekler et al. studies [[1]Hekler A. Utikal J.S. Enk A.H. Solass W. Schmitt M. Klode J. et al.Deep learning outperformed 11 pathologists in the classification of histopathological melanoma images.Eur J Cancer. 2019; 118: 91-96Abstract Full Text Full Text PDF PubMed Scopus (124) Google Scholar,[2]Hekler A. Utikal J.S. Enk A.H. Berking C. Klode J. Schadendorf D. et al.Pathologist-level classification of histopathological melanoma images with deep neural networks.Eur J Cancer. 2019; 115: 79-83Abstract Full Text Full Text PDF PubMed Scopus (106) Google Scholar]. Concordance rates estimated at a population level are higher [[3]Elmore J.G. Barnhill R.L. Elder D.E. Longton G.M. Pepe M.S. Reisch L.M. et al.Pathologists' diagnosis of invasive melanoma and melanocytic proliferations: observer accuracy and reproducibility study.BMJ. 2017; 357: j2813Crossref PubMed Scopus (224) Google Scholar,[4]Elder D.E. Piepkorn M.W. Barnhill R.L. Longton G.M. Nelson H.D. Knezevich S.R. et al.Pathologist characteristics associated with accuracy and reproducibility of melanocytic skin lesion interpretation.J Am Acad Dermatol. 2018; 79 (e5): 52-59Abstract Full Text Full Text PDF PubMed Scopus (19) Google Scholar]. Comparisons are problematic however, as existing studies apply multiple different diagnostic classes (3–5 or more), whereas despite claiming ‘pathologist-level classification’ [[2]Hekler A. Utikal J.S. Enk A.H. Berking C. Klode J. Schadendorf D. et al.Pathologist-level classification of histopathological melanoma images with deep neural networks.Eur J Cancer. 2019; 115: 79-83Abstract Full Text Full Text PDF PubMed Scopus (106) Google Scholar], the Hekler et al. studies are binary. It is nonetheless true that melanoma is currently a pathologist-defined entity, and lesions can be designated differently by different histopathologists. In the Hekler et al. studies, deep learning and the other histopathologists were in fact compared with a single pathologist's diagnostic view. A well-controlled study would require a large cohort of histopathology images containing melanomas and different nevi designated such by a number of experienced histopathologists (in our opinion at least 3, preferably more). Additionally, pathology and biology are not binary. Borderline lesions exist which cannot be unequivocally designated as melanoma or nevus [[5]Shain A.H. Yeh I. Kovalyshyn I. Sriharan A. Talevich E. Gagnon A. et al.The genetic evolution of melanoma from precursor lesions.N Engl J Med. 2015; 373: 1926-1936Crossref PubMed Scopus (629) Google Scholar]. Omitting such cases from the analysis and using binary choices instead of more sophisticated classifications such as the MPATH-DX (Melanocytic Pathology Assessment Tool and Hierarchy for Diagnosis) [[6]Lott J.P. Elmore J.G. Zhao G.A. Knezevich S.R. Frederick P.D. Reisch L.M. et al.Evaluation of the melanocytic pathology assessment tool and hierarchy for diagnosis (MPATH-Dx) classification scheme for diagnosis of cutaneous melanocytic neoplasms: results from the international melanoma pathology study group.J Am Acad Dermatol. 2016; 75: 356-363Abstract Full Text Full Text PDF PubMed Scopus (24) Google Scholar] or World Health Organisation [[7]Elder D.E. Massi D. Scolyer R.A. Willemze R. WHO classification of skin tumours.4th ed. IARC, Lyon2018Google Scholar] scheme may simplify study design and facilitate training of artificial intelligence. However, it does not reflect the current state of classification and makes application to unselected melanocytic lesions encountered in routine practice impossible. Addressed are the concerns we found most problematic, many of which were communicated to the Brinker group prior to publication (by us and others [personal communication]). We believe several criteria proposed to identify distortion or misinterpretation (‘spin’) in biomedical research studies [[8]Boutron I. Ravaud P. Misrepresentation and distortion of research in biomedical literature.Proc Natl Acad Sci U S A. 2018; 115: 2613-2619Crossref PubMed Scopus (114) Google Scholar,[9]Ochodo E.A. de Haan M.C. Reitsma J.B. Hooft L. Bossuyt P.M. Leeflang M.M. Overinterpretation and misreporting of diagnostic accuracy studies: evidence of “spin”.Radiology. 2013; 267: 581-588Crossref PubMed Scopus (113) Google Scholar] are present in the Hekler et al. studies. One wonders if most of the participating histopathologists chose not to be listed as an author or participant because of similar concerns regarding study design and/or the conclusions drawn. We do anticipate that computer-driven approaches including deep learning and artificial intelligence have the potential to revolutionise histopathology and provide an enormous diagnostic aid. However, the considerable hype regarding this topic and its potential should not alleviate the requirement to perform well-designed studies with critical scrutiny if the conclusions drawn are justified. Studies need to be performed on test sets with well-characterised tumour cohorts, image material suitable for histopathologic assessment and more adequate options for classification. A very fundamental flaw of the Hekler et al. study design is it impairs the diagnostic capabilities of the human pathologist, making a direct comparison with artificial intelligence problematic. Before entering patient care, computer-driven approaches will need to prove their value when compared with or supplementing human pathologists in an optimal diagnostic setting. The authors have no conflicts of interest to declare. No third-party funding was applied to the current manuscript. Funding agencies did not influence the content of the manuscript. Deep learning outperformed 11 pathologists in the classification of histopathological melanoma imagesEuropean Journal of CancerVol. 118PreviewWith limited image information available, a CNN was able to outperform 11 histopathologists in the classification of histopathological melanoma images and thus shows promise to assist human melanoma diagnoses. Full-Text PDF Open AccessReply to the letter to the editor: ‘Deep learning outperformed 11 pathologists in the classification of histopathological melanoma images’European Journal of CancerVol. 130PreviewAt the end of the recruitment period, 14 addresses had acknowledged the invitation by e-mail, and these physicians were counted as invited to the survey. As described in the manuscript, participants with and without board certification were eligible because pathologists with and without board certification assess histologic slides in clinical routine. Full-Text PDF Open Access

Read the paper · More papers on PaperTik