Optimization of genomic classifiers for clinical deployment: evaluation of Bayesian optimization to select predictive models of acute infection and in-hospital mortality

Michael B. Mayhew, Elizabeth Tran, Kirindi Choi, Uros Midic, Roland Luethy, Nandita Damaraju, Ljubomir J. Buturović · 2020

Patient lives depend on rapid and accurate detection of acute infection status as well as prediction of the severity of their condition. However, conventional clinical measures for these tasks (e.g. physiological symptoms, lab culture results) are suboptimal. Characterizing a patient's immune response from gene expression measurements derived from blood has proven an effective strategy for timely detection of the presence, type and severity of infection in that patient. Developing machine learning classifiers from these gene expression measurements depends, in part, on hyperparameter optimization (HO), a process by which aspects of the classifier related to its training or structure are tuned according to some performance objective. Multiple approaches including grid search, random sampling, and Bayesian optimization have been proposed and proven successful for HO of machine learning classifiers. In general, these approaches have been compared under a limited range of settings and solely with respect to the internal validation performance of the tuned classifier (e.g. in k-fold cross-validation or on a validation set partitioned from the training set). In addition, these approaches have not been evaluated previously for HO of diagnostic classifiers using genomic data. In this comprehensive analysis, we evaluate performance in both internal validation and in a multi-cohort, held-out external validation dataset of multiple types of classifiers selected by three methods for HO (including Bayesian optimization). We evaluate performance on two different classification tasks: 1) multi-class acute infection detection ( BVN ; B acterial infection, V iral infection or N oninfected inflammation) and 2) prediction of a mortality event within 30 days of hospital admission. We train and evaluate all classifiers with multi-cohort expression datasets comprising a variety of geographical regions, healthcare settings, and assay platforms. We find that classifiers selected by all three hyperparameter optimization approaches perform comparably but that multi-layer perceptron classifiers of 30-day mortality selected by Bayesian optimization outperform their counterparts selected by other HO approaches. Contrary to previous research, we find evidence in our setting that: 1) Bayesian optimization is not necessarily more efficient in selecting classifiers compared with alternative HO methods and 2) using a common variant of Gaussian process-based Bayesian optimization (i.e. automatic relevance determination) results in only marginal gains in classifier performance.

Read the paper · More papers on PaperTik