A Scalable Feature Selection and Model Updating Approach for Big Data Machine Learning

Baijian Yang, Tonglin Zhang · 2016

In this paper, we proposed an innovative approach for feature selection and model updating in big data machine learning. Since hard drive access is the biggest barrier for big data problems, it is therefore nature to reduce disk I/O operations when evaluating different combinations of features, or updating a learning machine. Particularly, we are interested in discovering if small enough matrices exist to represent a system and if the calculation of such matrices can be achieved in a row-by-row fashion to avoid read data from hard drive over and over again. We examined the case of linear regression and proved that arrays of sufficient statistics can be used for feature selection and model updating. Algorithms were designed to compute the arrays in both single processor and MapReduce fashion. The proposed approach can reduce the memory requirement down to O(p2), where p is the number of variables in the data set. Simulation results also demonstrated the effectiveness of the algorithms with major computation improvements.

Read the paper · More papers on PaperTik