Variant relevance prediction in extremely imbalanced training sets

Max Schubach, Matteo Ré, Peter N. Robinson, Giorgio Valentini · 2017

The interpretation of non-coding variants still constitutes a major challenge in the application of whole-genome sequencing. For example for Mendelian diseases only several hundreds of pathogenic regulatory mutations are known but millions of possible neutral sites can be derived. In this context, machine learning (ML) methods for predicting disease-associated non-coding variants are faced with a chicken and egg problem - such variants cannot be easily found without ML, but ML cannot be applied effectively until a sufficient number of instances has been found. Recent ML-based methods for variant prediction do not adopt specific imbalance-aware learning techniques to deal with imbalanced data that naturally arise in several genome-wide variant scoring problems, resulting in relatively poor performance with reduced sensitivity and precision. Here, we present a ML algorithm based on resampling techniques and a hyper-ensemble approach, called hyper SMOTE Undersampling with Random Forests (hyperSMURF), which is able to deal with extremely imbalanced datasets. HyperSMURF outperforms previous methods on two different published imbalanced variant datasets: regulatory Mendelian mutations and classification of microRNA/SNP pairs into eQTLs or non-eQTLs. We show that imbalance-aware ML is a key issue for the design of robust and accurate prediction algorithms and the provided method hyperSMURF can be applied effectively to discover disease-associated variants out of millions of neutral sites from whole-genome sequencing.

Read the paper · More papers on PaperTik