Deep Learning Makes Its Way to the Clinical Laboratory
Ronald Jackups · Clinical Chemistry · 2017
Are pathologists obsolete? Although the current answer is most certainly no, this question will be asked with increasing intensity as technological advances in clinical diagnostics, such as machine learning and artificial intelligence, prove to be as effective or even more accurate than the judgments of human experts. One area of pathology experiencing rapid expansion in the diagnostic power of computer technology is automated digital image analysis (DIA)2 (1). A recent study has suggested that automated scoring of biomarkers in breast cancer (Ki67, ER, PR, and HER2) by DIA can predict a molecular subtype of cancer with higher sensitivity and specificity than manual scoring by board-certified pathologists (2). The prospect of using such a tool is not restricted to research, as analyzers have been approved by the Food and Drug Administration (FDA) to score these biomarkers (3, 4). Clinical pathology has also become a fertile ground for DIA in the form of an FDA-approved analyzer for peripheral blood smear review (5). This analyzer identifies, classifies, and quantifies leukocytes by morphologic type using a machine learning approach to assist the technologist or pathologist performing manual differential counts. However, there currently is no similar tool for the morphologic characterization of erythrocytes (RBCs) that has been approved by the FDA. In this issue of Clinical Chemistry, Durant et al. present a novel DIA tool that uses very deep convolutional neural networks (CNNs) to predict the morphology of individual RBCs for the identification of significant abnormal populations that may aid in the diagnosis of hematologic and nonhematologic disease (6). They find that their method outperforms a non–FDA-approved analyzer similar to that used to characterize leukocyte morphology (7, 8). Their method offers not only high accuracy in morphologic profiling but also could provide higher consistency, reproducibility, and efficiency over the current slow, manual process of peripheral smear review. The algorithm used by Durant et al. is an example of deep learning, a relatively recent and powerful set of machine learning techniques designed to classify new objects, such as RBCs, using algorithms “learned” from previous patterns (9). Deep learning has been used in many applications both within and outside of medicine, including image recognition, speech recognition, drug discovery, and mutation analysis (9). Although the study by Durant et al. represents the first use of deep learning for DIA in a clinical pathology application, it has already been explored for use in anatomic pathology (10, 11). Deep learning differs from earlier machine learning techniques in several ways (9–11), most notably its ability to learn “features,” which are informative components of the objects to be classified. Whereas less sophisticated algorithms must rely on a limited number of preselected features (e.g., central pallor and membrane spicules in RBCs), deep learning identifies its own most useful set of features. Another advantage offered by the CNN used by Durant et al. is the ability to operate on small local fields of an image rather than on the entire image at once; this allows the model to ignore the orientation or rotation of an image, which would otherwise be difficult to account for in a peripheral smear in which RBCs are scattered haphazardly (6). Finally, as its name implies, another benefit of deep learning is its depth, or ability to pool information from many layers of input from large amounts of training data to improve classification accuracy. However, deep learning, as well as machine learning in general, has its own limitations and challenges and has faced criticism from many angles (11, 12). Perhaps the most apparent limitation is that it is poorly understood by laypeople, including most physicians. It represents a “black box,” such that when a classification outcome is unexpected or unrealistic, it may be difficult to identify the source of error or explain the process in an intuitive way. One of the advantages of deep learning (unsupervised learning of features) becomes a handicap when it is necessary to explain what features the algorithms are using to make predictions. In addition, the field of machine learning uses terminology that may confuse physicians, particularly laboratorians, such as “recall” instead of “sensitivity” and “precision” instead of “positive predictive value.” Although these terms help to highlight the practical, functional, and inferential differences between machine learning on large data sets and simple measures of test accuracy, they obscure the fact that, at its core, machine learning is a diagnostic tool that operates in much the same way as any diagnostic laboratory test. As machine learning continues to grow in popularity, it will continue to elicit criticism of its limitations. Although some of this criticism is because of its opaqueness, it is important to understand how and why it can fail to perform its stated task (12). Errors may be because of how the training set was constructed. If the experts who label the training set with “correct” classifications make mistakes or even disagree on the proper labels, the algorithm may propagate those errors. Fortunately, this may be mitigated if most of the labels are correct, as the algorithm will minimize the effect of incorrect labels. Durant et al. discovered this when they noted that a few of the RBCs in their test set had been incorrectly labeled, leading their algorithm to make the “wrong” but clinically appropriate prediction (6). Deep learning methods may also be sensitive to the size of the data available, requiring large training sets to avoid overfitting errors. Durant et al. also noted this, as the classification accuracy was much higher for RBC types that were well-represented in their training set (e.g., schistocytes and target cells) than for RBC types that were poorly represented (e.g., dacryocytes). Lastly, the most insidious limitation of machine learning is its inability to incorporate information and knowledge not included in the original training set. In medicine, this means that qualitative and complex patient characteristics (e.g., patient history) do not contribute to classification (12). Clinical context is critical in reviewing DIA results, as abnormal cells may reasonably be interpreted differently based on patient diagnosis. In reality, this is no different from any single laboratory test, and all current FDA-approved DIA devices can serve only as an aid to the operator or pathologist, who is responsible for the final classification (3–5). Ultimately, however, it may be possible to provide enough input data to an algorithm like that of Durant et al. for it to act as a primary diagnostic device. Despite these limitations, there are many opportunities for incorporating machine learning into laboratory medicine (13). With so much discrete data stored for each patient, laboratory information systems can provide a wealth of training input to lead to highly reliable diagnoses. In addition, images from DIA outputs could flow into the electronic medical record to assist clinicians with treatment decisions. There is great potential for technological growth if there are pathologists trained in informatics who can design and implement these tools effectively. Are pathologists obsolete? No, and they never will be if they remain on the forefront of new technologies like machine learning, both in development and integration. digital image analysis Food and Drug Administration erythrocytes convolutional neural network.