Feature Selection for Big Data Based on Mapreduce and Voting Mechanism
Sufang Zhang, Junhai Zhai, Tian Shi, Xiang Zhou, Li Yan · 2020
With the rapid development of computer network technology and wireless sensor technology, as well as the arrival of the era of big data, the dimension and sample number of data are growing rapidly. Accordingly, it is important to investigate the problem of feature selection for big data and to design feature selection algorithm for big data. Based on MapReduce and voting mechanism, a feature selection method for big data is proposed in this paper. The proposed methods include three steps: Firstly, partition big data set into m subsets, and deploy the subsets to m computing nodes of Hadoop. Secondly, on the m computing nodes, we employ a feature selection algorithm based on genetic algorithm to select important features in parallel using local data subset, and obtain m feature subsets. Finally, for each feature, m feature subsets are used to vote on it, and the final feature subset is selected according to the voting results. Experimental results on four big data sets demonstrate that the proposed method is effective and efficient.