Building an SVM Classifier for Automated Selection of Big Data

Junhua Ding, Jiabin Wang, Xiaojun Kang, Xin‐Hua Hu · 2017

The quality of big data could great impact the value extracted from the data. Automated filtering of noisy data from big data is an ideal approach for improving the quality of big data. However, due to large volume and variety of big data, automated filtering of noisy data from big data is a grand challenging task. In this paper, we propose a support vector machine based approach for automated classification of big data so that the noisy data are classified as separated categories from the regular data. In order to improve the classification accuracy and training performance, we design an experiment for improving the classification model through finding the optimized learning feature set and an approach for iteratively improving the quality of the training data set. We conducted a thorough experimental study of automated classification of massive image data of biology cells to explain the approach of automated selection of big data and demonstrate its effectiveness. Finally, we compare the performance of the SVM based classification and a deep learning based classification of the same data set. The proposed approach and experience collected from the experimental study can help big data researchers and practitioners to design strategies for improving the quality of big data, designing high performance classifier, and building tools for automated selection of big data.

Read the paper · More papers on PaperTik