Robust feature selection with penalized regression in imbalanced high dimensional data
Jie Ren · University of Southern California Digital Library · 2014
This work is motivated by an ongoing USC/Illumina study of prostate cancer recurrence after radical prostatectomy. The study generated gene expression data for nearly thirty thousand probes from 187 tumor samples, of which 33 came from patients with recurrent prostate cancer and 154 came from patients with non?recurrent prostate cancer after years of follow?up. Our goal was to use penalized logistic regression and stability selection to find a ?gene signature? of recurrence that can improve upon PSA and Gleason?score, which are well?established but poor predictors. For interpretability and future clinical use, the gene signature should ideally involve a small proportion of probes to predict recurrence in new patients. Due to the skewness in the data, the model selected by tuning the LASSO penalty parameter based on the average misclassification rate in cross validation did not have a balanced performance, i.e. it predicted non?recurrent cancer with high accuracy but predicted recurrent cancer with very low accuracy. In addition, standard penalized regression with cross validation appeared to select many noise features. In my simulation study in Chapter 2, I evaluated the performance of models selected by different metrics in imbalanced data. I concluded that G?mean?thr (G?mean with an alternative cutoff) and area under the ROC curve (AUC?ROC) were the most robust metrics to class imbalance. In Chapter 3, I examined the performance of stability selection (SS?thr) in simulation studies and found that its feature selection capability (a) depended on the stability cutoff chosen and (b) is conservative as a result of a stringent error control. To address these problems, I proposed new feature selectors based on stability selection, including SS?test, an essentially parameter?free test?based outlier?detection approach, and SS?rank and SS?ranktest, parameter?free rank?based methods. I demonstrated their advantage over SS?thr, ULR, and LASSO with cross validation in extensive simulation studies and also found that all these stability selection based methods were robust to class imbalance. These newly developed methods and procedures was applied to the prostate cancer recurrence data. I used a variety of metrics to do model selection within the penalized logistic regression framework using imbalanced prostate cancer recurrence data and demonstrated that G?mean with the case?proportion cutoff selected the model with the most balanced prediction accuracy in cross validation. In addition, I also showed the importance of using an appropriate cutoff to evaluate models when models were built from skewed data. I also applied stability selection based methods including SS?thr, SS?test, SS?rank and SS?ranktest to select important genes from the same data. Three genes, ABCC1, NKX2?1 and ZYG11A were identified by all of the methods and also appeared to stand out from other features in the stability path plot. I fit a logistic regression model using these genes and clinical features, which has significantly higher prediction accuracy than clinical?only models when evaluated by cross?validation.