Predicting Implantation Outcome from Imbalanced IVF Dataset
Asli Uyar, Ayşe Bener, Haydar Nadir Ciray, Mustafa Bahçeci · 2009
Predicting implantation outcomes of in- vitro fertilization (IVF) embryos is critical for the suc- cess of the treatment. We have applied Naive Bayes classifler to an original IVF dataset in order to dis- criminate embryos according to implantation poten- tials. The dataset we analyzed represents an imbal- anced distribution of positive and negative instances. In order to deal with the problem of imbalance, we ex- amined the efiects of over sampling the minority class, under sampling the majority class and adjustment of the decision threshold on the classiflcation perfor- mance. We have used features of Receiver Operat- ing Characteristics (ROC) curves in the evaluation of experiments. Our results revealed that it is possi- ble to obtain optimum True Positive and False Pos- itive Rates simply by adjusting the decision thresh- old. Under-sampling experiments show that we can achieve same prediction performance with less data as well as 736 embryo samples. samples with positive outcomes. Any classifler built on these datasets has much more information to identify un- successful IVF treatments compared to successful ones. Therefore, implantation prediction is handled as a typical case of learning from imbalanced data problem. We ana- lyze the efiects of re-sampling the training data and de- cision threshold optimization on imbalanced IVF dataset using Naive Bayes classifler. Our results show that 0.3 is the best threshold for classiflcation of embryos. We have also considered another research problem that is the determination of the smallest amount of training data required to build an efiective predictor model. Data collection is a costly and time-consuming process in medi- cal applications. Analysis of under-sampling experiments leaded to deflne su-cient size of embryo samples for im- plantation prediction that would reduce the efiort spent for data collection in IVF domain.