Machine Learning for Prediction in Electronic Health Data
Sherri Rose · JAMA Network Open · 2018
Machine learning for prediction in electronic health data has been deployed for many clinical questions during the last decade.Machine learning methods may excel at finding new features or nonlinear relationships in the data, as well as handling settings with more predictor variables than observations.However, the usefulness of both these data and machine learning has varied.Electronic health data often have quality issues (eg, missingness, misclassification, measurement Of course, assessing the generalizability of a prediction algorithm goes well beyond using crossvalidated metrics to evaluate overfitting.Wong et al 1 carefully discussed a number of limitations in their work, including the lack of external validation in other health systems.As in similar studies, the step from good performance in the study to a generalizable algorithm is vast and sometimes may not be feasible.Patients receiving treatment in varied care settings or geographic regions simply may require tailored tools.Recognizing the need for unique tools in different populations is not inherently negative, but one of many considerations not magically solved by using machine learning.Even those algorithms that prove to be generalizable may quickly become outdated as treatment patterns or physician incentives to code health conditions change.4 Increased social tolerance for certain +