Semi-supervised Multi-Label Feature Selection Using Hessian Energy based on Maximum Relevance and Minimum Redundancy

Xinping Wu, Hongmei Chen, Tianrui Li, Hao Chen, Chuan Luo · 2021 16th International Conference on Intelligent Systems and Knowledge Engineering (ISKE) · 2021

With the development of the Internet, instances are not limited to a single label. “High-dimensional and multiple labels” poses a huge challenge to the use of data information. Multi-label feature selection technology has attracted much attention as one of the important dimensionality reduction methods. The labeling cost of the data is too large, which leads to the widespread application of semi-supervised multi-label feature selection in feature selection. However, the existing semi-supervised multi-label sparse feature selection algorithm considers only the correlation between features and labels, without taking into account redundancy among features; moreover, most of them capture the local structure of the data by using Laplacian regularization. But the Laplacian regularization method is not good at inferring the popular structure. Therefore, a novel Semi-supervised Multi-Label Feature Selection algorithm using Hessian Energy based on maximum relevance and minimum redundancy (S2MFSHMRMR) is proposed in this study. Firstly, Hessian regularization and Hilbert-Schmidt independence criterion (HSIC) are respectively used to reflect the inherent local geometric characteristics of the data and to solve the problem of the relevance between features and multiple labels. Then, a new custom redundancy regularization method is proposed, and the redundancy between features is measured by GMM-based Bhattacharyya distance. In addition, the soft labels required for the Bhattacharyya distance are obtained by using the multi-label label propagation method. Finally, a closed-form method based on the matrix Lagrangian multiplier method is proposed to minimize the objective function. Experiments are performed to verify the effectiveness of the method. The proposed algorithm S2MFSHMRMR is compared with five algorithms on six Mulan data sets. Experimental results show that the proposed method is very competitive.

Read the paper · More papers on PaperTik