MLCR: A Fast Multi-label Feature Selection Method Based on K-means and L2-norm

Amin Hashemi, Mohammad Bagher Dowlatshahi · 2020

Feature selection is an essential step in data mining and machine learning that increases classification accuracy and reduces the computational time by eliminating redundant and unrelated features. In this paper, a fast feature selection algorithm is introduced based on clustering ranking in feature-label space and L2-norm, called MLCR. This method is a filter-based method for multi-label datasets. We used a two-step strategy for this method. First, we used the k-means algorithm to cluster the features based on their correlation with labels. Then we sorted the features in each cluster based on L2-norm in descending order and finally set rank to each feature. This will allow similar features to be grouped into one cluster. In the second step, the features with the same rank are sorted like the previous step and added to the feature ranking vector. To verify the efficiency of MLCR, we have compared the obtained results of this method with five well-known multi-label feature selection algorithms based on various real-world multilabel datasets in different dimensions. The results demonstrate that our proposed method outperforms the other methods in the classification measures and run-time.

Read the paper · More papers on PaperTik