Commentary: Towards machine learning-enabled epidemiology
Louisa R. Jorm · International Journal of Epidemiology · 2020
Since first emerging as a discipline in the 1990s, data science has become a critical area of workforce skills shortage.1 Although data science has no agreed definition, it is centred in multidisciplinary and interdisciplinary approaches to extracting knowledge or insights from data for use in a broad range of applications.1 The role of the epidemiologist in the health and medical domain aligns strongly with a common definition of a data scientist as someone who ‘combines domain-specific expertise with analytic skills to extract knowledge from data to drive action’.2 However, most training programmes in epidemiology do not teach the primary skills that healthcare organizations seek in data scientists, which include machine learning (ML) and the open-source programming languages R and Python.3 Indeed, a course in data science was a mandatory component of only 18% of epidemiology programmes offered by the top 20-ranked public health schools in the USA in 2019.4 There has been considerable discussion within the statistical community regarding the relationship between statistics, data science and ML,5 emphasising the need to ensure that statisticians have the necessary skills in computation. Engineering and computer science graduates are seen as currently better equipped than statisticians to contribute as data scientists.6 Forging new approaches that bring together ML and statistical communities and mindsets is presented as a solution to addressing challenges inherent in the application of ML to big datasets including selection bias, measurement error, quantifying uncertainty, and interpretability.7 It is still early days for similar discussions among epidemiologists. However, commentators argue that whereas epidemiologists do not necessarily need to learn coding at the expense of core epidemiological skills,4,8 or become experts in ML,4 they do need a foundational knowledge of data science techniques to equip them to work in the large interdisciplinary teams that will make big discoveries in science. The pervasive use of closed-source programming languages (e.g. SAS, Stata) is cited as being a barrier to the integration of ML techniques in epidemiology.9 The burgeoning use of ML across all aspect of health and medicine creates an imperative for epidemiologists to be at the very least ‘ML-aware’. Three papers in this issue of the International Journal of Epidemiology serve to advance this cause. All three focus on the unique power of ML methods for prediction. Blakely et al.10 explore how supervised ML methods can be used as an approach to ‘best prediction’ in the underpinning steps in contemporary causal inference methods, and touch briefly on their potential for robust detection and estimation of heterogeneity of treatment effects. Examples of the latter application of ML in health and medicine are proliferating, using both secondary analysis of randomized controlled trials and large observational datasets. Bannick et al.11 present a framework for constructing ensemble models—which combine multiple ML algorithms to improve predictive performance—for descriptive epidemiology and describe applications in burden of disease estimation. Importantly, they have made an interactive example of their Cause of Death Ensemble Model (CODEm) available using Jupyter Notebook (an open-source web application for creation and sharing of documents that contain live code, equations, visualizations and text) and GitHub (an online code hosting platform for collaboration and version control). These essential data science tools serve both to promote reproducible science and provide examples for training and education, but unfortunately are only rarely used by epidemiologists to share their work. Rose12 provides an overview of the penetration of ML methods into health services research for applications including prediction of health care spending and clinical quality, as well as causal effect estimation in comparative effectiveness research and policy evaluation. Unique among the three papers, she touches on the potential uses of unsupervised ML methods. In fact, the definition of ML given in the glossaries of both Blakely et al.10 and Bannick et al.11 is specific to supervised ML: ‘Algorithms that aim to “learn” or predict outputs from inputs (covariates) based on a dataset that contains both inputs and labelled output’. Box 1 provides a frequently cited, more general definition of ML as well as definitions and examples for four commonly described subtypes: supervised ML, unsupervised ML, reinforcement learning and deep learning. Definitions of machine learning and subtypes Machine learning (adapted from13,14) A family of mathematical modelling techniques that uses a variety of approaches to automatically learn from data, without explicit programming. Supervised machine learning (adapted from15,16) Focuses on learning from a collection of labelled examples. Each example (e.g. patient) is represented by input data (e.g. demographics, vital signs, laboratory results) and a target label (such as being diabetic or not). The learning algorithm then seeks to learn a mapping from the inputs to the labels that can generalize to new examples. When target variables are continuous real number values, the supervised learning task(s) are known as regression problems, and when the target variables are categorical variables, the tasks are known as classification problems. Common supervised learning algorithms include linear regression, logistic regression, decision trees, random forests, support vector machines (SVM) and k‐nearest neighbours. Unsupervised machine learning (adapted from15,16) Reinforcement learning (adapted from17) Machine learning (adapted from13,14) A family of mathematical modelling techniques that uses a variety of approaches to automatically learn from data, without explicit programming. Supervised machine learning (adapted from15,16) Focuses on learning from a collection of labelled examples. Each example (e.g. patient) is represented by input data (e.g. demographics, vital signs, laboratory results) and a target label (such as being diabetic or not). The learning algorithm then seeks to learn a mapping from the inputs to the labels that can generalize to new examples. When target variables are continuous real number values, the supervised learning task(s) are known as regression problems, and when the target variables are categorical variables, the tasks are known as classification problems. Common supervised learning algorithms include linear regression, logistic regression, decision trees, random forests, support vector machines (SVM) and k‐nearest neighbours. Unsupervised machine learning (adapted from15,16) Reinforcement learning (adapted from17) Definitions of machine learning and subtypes Machine learning (adapted from13,14) A family of mathematical modelling techniques that uses a variety of approaches to automatically learn from data, without explicit programming. Supervised machine learning (adapted from15,16) Focuses on learning from a collection of labelled examples. Each example (e.g. patient) is represented by input data (e.g. demographics, vital signs, laboratory results) and a target label (such as being diabetic or not). The learning algorithm then seeks to learn a mapping from the inputs to the labels that can generalize to new examples. When target variables are continuous real number values, the supervised learning task(s) are known as regression problems, and when the target variables are categorical variables, the tasks are known as classification problems. Common supervised learning algorithms include linear regression, logistic regression, decision trees, random forests, support vector machines (SVM) and k‐nearest neighbours. Unsupervised machine learning (adapted from15,16) Reinforcement learning (adapted from17) Machine learning (adapted from13,14) A family of mathematical modelling techniques that uses a variety of approaches to automatically learn from data, without explicit programming. Supervised machine learning (adapted from15,16) Focuses on learning from a collection of labelled examples. Each example (e.g. patient) is represented by input data (e.g. demographics, vital signs, laboratory results) and a target label (such as being diabetic or not). The learning algorithm then seeks to learn a mapping from the inputs to the labels that can generalize to new examples. When target variables are continuous real number values, the supervised learning task(s) are known as regression problems, and when the target variables are categorical variables, the tasks are known as classification problems. Common supervised learning algorithms include linear regression, logistic regression, decision trees, random forests, support vector machines (SVM) and k‐nearest neighbours. Unsupervised machine learning (adapted from15,16) Reinforcement learning (adapted from17) Rose12 touches on the uses of ML for clustering and dimensionality reduction and Blakely et al.10 mention briefly its potential uses for addressing measurement error and missing data. However, none of the three current papers explore how these methods can expand the universe of data available for epidemiological research through enabling use of the very large observational datasets that are increasingly available through electronic medical records (EMRs), as well as emerging data streams from sensors and wearables. A typical EMR consists of multimodal data including unstructured text (e.g. clinical notes, discharge summaries), 2D and 3D images (e.g. X-ray, magnetic resonance imaging), time-series signals (e.g. electrocardiogram traces) and PDF files with text (e.g. lab reports). EMR data present significant analytic challenges because of the very large numbers of variables, the sparse nature of much of the data, the irregular time series and the ‘messy’ nature of free text data. The complexities of EMRs compound the already time-consuming tasks of data preparation and data cleaning, which already absorb between 60 and 80% of time taken in analysis tasks.20 Recently, a variety of ML methods for interactive data cleaning have been described, which are successful in tasks from removal of duplicates from coded clinical data21 and resolving data inconsistencies in EMRs22 through to interactive iterative data cleaning using a sophisticated algorithm that prioritizes for cleaning those records that are most likely to change model predictions.23 Privacy protection through de-identification of sensitive variables is another key task in preparing observational datasets, the difficulty of which is vastly increased when dealing with unstructured text that contains patient and provider names and other identifiers. Named entity recognition (NER), a subtask of natural language processing (NLP), classifies entities from unstructured text into pre-defined categories such as persons, locations and organizations. State-of-the-art NER systems using ML produce near- or better-than-human performance, and recently recurrent neural networks have demonstrated excellent performance in de-identifying EMR text data.24,25 ML-based algorithms for NLP have also demonstrated excellent performance, compared with manual abstraction, for automatic coding of clinical notes according to diagnosis or disease codes (e.g.26–28), with deep learning outperforming other ML techniques.29 ML methods also offer great promise for handling missing values, a major problem in EMRs, as well as many other datasets used in epidemiology. Recently, generative adversarial networks (GANs), which pit two neural networks against each other, have been demonstrated to significantly outperform state-of-the art imputation techniques in medical data (including multivariate imputation by chained equations [MICE]).30,31 Finally, ML methods applied to EMRs can be used to identify patients with specific characteristics of interest (either exposures or outcomes), a process known as electronic phenotyping,32 which can extend beyond simple identification of patients with well-defined diagnoses or risk factors to discovery of subtypes of patients with complex conditions based on multiple features. Phenotyping may be used to support a range of epidemiological study designs, including cohort and case–control analyses using routinely collected data. Phenotypes can be incorporated as predictors into classical epidemiological analyses, and their identification represents the first step in precision approaches to disease prevention and management. Unsupervised clustering methods, supervised methods including support vector machines, and joint use of supervised and unsupervised approaches have been employed,32,33 with recent work demonstrating superior performance of deep learning methods over other NLP methods for phenotyping using clinical text.32 To ensure that epidemiologists take their place as health and medical domain experts in data science, priorities for skills and methods development include the following: Understanding the foundational concepts of ML and types of ML models. Familiarity with ML terminology, which often differs from that used in epidemiology and statistics for the same concepts (a useful resource is here34). Proficiency in coding in R and Python. Sharing of code and use of data science tools for reproducible research including GitHub and Jupyter Notebook. Finding and appraising scientific literature from the domains of computer science and engineering, which often takes the form of conference papers and is equation-heavy. Using ML to enhance classical epidemiology through improving descriptive epidemiology, outcome and exposure measurement and causal estimation. Engaging with ML practitioners to develop analytic approaches for EMRs and other big data that incorporate epidemiological thinking, including robust consideration of confounding, causality and bias. Advancement of epidemiology as an ML-enabled discipline will help to drive the appropriate use of ML to improve health, healthcare and health equity, and avoid potential harms inherent in the uncritical uptake of black box ‘AI’ (artificial intelligence) solutions in health care. None declared.