A feature selection algorithm for multilayer perceptron based on simultaneous two-sample representation

Shudong Liu, Ke Zhang, Xu Chen · 2020

Classification is one of the hot topics of machine learning domains, its main task is to learn a classification model from training data and predict the labels of unknown samples. To date, many classification models have been proposed and are widely used in various realworld applications, e.g., naive Bayes (NB), logistic regression (LR), support vector machine (SVM) have been successfully employed in spam recognition, bank loan credit scoring and network rumor recognition, respectively. Imbalance learning is an important branch of classification task in machine learning domains. Data-level, algorithm-level and ensemble solutions are the three main methods proposed thus far to address imbalance learning. To alleviate the issues of data explosion and feature selection for a multilayer perceptron with simultaneous two-sample representation, in this paper, we propose a novel feature selection method based on the pairwise samples distance constraint, which considers the class labels of paired samples, select the features which push two similar samples closer together and pull two different samples farther apart. Finally, we conduct experiments on four high-dimensional DNA microarray datasets. The experimental results demonstrate that our proposed algorithms outperform some state-of-the-art algorithms in terms of F-measure and G-mean.

Read the paper · More papers on PaperTik