Machine Learning for Biomarker Discovery in Cancer Pharmacogenomics Data
Arvind Singh Mer, Petr Smirnov, Benjamin Haibe‐Kains · 2019
Over the past decade, there has been an explosion in the availability of massive datasets combining drug screening with high-throughput molecular profiling in cancer model systems. These datasets have become a rich community resource which can be leveraged for biomarker discovery, in-silico validation, drug repurposing, drug method of action prediction, and to train statistical machine learning models for drug response prediction. However, this data poses unique challenges during analysis and requires methods that are robust to the noise inherent in the drug sensitivity assays. Furthermore, irreproducibility of some findings across studies strongly motivates integrative analysis across studies. Fortunately, tools have been developed implementing bioinformatics and machine learning methods designed specifically for the analysis of pre-clinical pharmacogenomics data. In this tutorial, participants will become familiar with common preclinical cancer models (such as cell-line, patient derived xenografts and organoids) and publicly available large pharmacogenomics datasets. In the hands on session, they will be introduced to the tools and packages published for analysis of these datasets, with a focus on tools written in R. Furthermore, after becoming familiar with the challenges posed by the noise in the pharmacological assays observed in high-throughput pharmacogenomics, participants will gain hands on experience using these datasets for the purpose of biomarker discovery and validation as well as building machine learning models predictive of drug response. The focus will be on translational research, validating discoveries from in vitro datasets using in vivo pharmacogenomic and clinical datasets.