Matthews correlation coefficient-based feature ranking in recursive ensemble feature selection for high-dimensional and low-sample size data
David Rojas-Velázquez, Aletta D. Kraneveld, Alberto Paolo Tonda, Alejandro Lopez‐Rincon · Machine Learning with Applications · 2025
High-dimensional, low-sample-size (HDLSS) omics datasets pose significant challenges for biomarker discovery due to class imbalance and reproducibility issues. We present MCC-REFS, an enhanced version of the Recursive Ensemble Feature Selection (REFS) method, which incorporates the Matthews Correlation Coefficient (MCC) as a feature selection metric to improve robustness and performance. MCC-REFS leverages eight machine learning classifiers in an ensemble framework and does not require predefined feature counts or complex hyperparameter tuning. We evaluated MCC-REFS against REFS, GRACES (GRAph Convolutional nEtwork feature Selector), DNP (Deep Neural Pursuit), and GCNN (Graph Convolutional Neural Network) using synthetic datasets and real-world mRNA datasets (Colon, SMK-CAN-187, ALLAML) as well as a multi-label breast cancer dataset from TCGA. MCC-REFS consistently selected more compact and informative feature sets, achieving higher or comparable classification performance. Notably, in the TCGA dataset, MCC-REFS selected 327 genes and achieved an accuracy of 0.9602, outperforming GCNN (0.9133). The method demonstrated superior adaptability to class imbalance and flexibility in feature selection, validated using independent classifiers such as MLP (Multilayer Perceptron). These results suggest MCC-REFS is a robust and scalable tool for feature selection in omics-based diagnostic and prognostic applications. • Tested against deep learning methods using real-world omics datasets. • Outperformed other methods on binary-label and multi-label datasets. • Does not require fixing the number of target features. • Suitable for unbalanced class datasets.