Reply: metrics to assess machine learning models

Alvin Rajkomar, Andrew M. Dai, Mimi Sun, Michaela Hardt, Kai Chen, Kathryn Rough, Jeffrey Dean · npj Digital Medicine · 2018

We thank Prof. Pinker for bringing up important points on how to assess the performance of machine learning models. The central finding of our work is that a machine learning pipeline operating on an open-source data-format for electronic health records can render accurate predictions across multiple tasks in a way that works for multiple health systems. To demonstrate this, we selected three commonly used binary prediction tasks, inpatient mortality, 30-day unplanned readmission, and length of stay, as well as the task of predicting every discharge diagnosis. The main metric we used for the binary predictions was the area-under-the-receiver-operator curve (AUROC).

Read the paper · More papers on PaperTik