A Tabular Variational Auto Encoder-Based Hybrid Model for Imbalanced Data Classification With Feature Selection

Asha Abraham, Habeeb Shaik Mohideen, Ramalingam Kayalvizhi · IEEE Access · 2023

Cancer is the deadliest disease in humankind. Ovarian Cancer (OC) stands important among female-specific cancers. Epithelial Ovarian Cancer (EOC) is the most commonly occurring subtype of OC. The disease is identified in later stages due to the unrevealed symptoms in the early stages. Gene Expression experiments along with machine learning (ML) methodologies can lead to preventive care of OC by early identifying the malignant gene transformations and to use precision medicine which aids in fast recovery. The proposed hybrid Tabular Variational Auto Encoder oriented dictionary based Stratified K Fold Cross Validation (TVAE_dict_SKCV) is an effective model to handle the threat. The main objective is to assess the significance of EOC screening variables for categorizing high-risk patients. It initially generated synthetic data using the TVAE model to increaseF the EOC subtype data size taken from the Cancer Cell Line Encyclopedia (CCLE). The synthesized data were balanced utilizing the Synthetic Minority Oversampling Technique (SMOTE). Significant features got selected with Boruta Feature Selection (FS) method. The HYPERPARAMETERS were fine-tuned employing Optuna optimizer and applied enhanced SKCV with Random Forest (RF) classifier. The TVAE_dict_SKCV method with Boruta acquired an accuracy of 98.5 % which outperforming the experiment with Lasso FS and with original data. The Pickle tool reserved the concealed parameter values of the model. Shapley Additive explanations (SHAP) summarize the main features which classify. Optuna efficiently reduced the computing time than the Grid search CV (GsCV) optimizer.

Read the paper · More papers on PaperTik