Nature Inspired Instance Selection Techniques for Support Vector Machine Speed Optimization

Andronicus A. Akinyelu, Absalom El-Shamir Ezugwu · IEEE Access · 2019

Due to the fast-growing rate of information sources, many organizations and individuals are overwhelmed with vast amounts of data. The rate of data growth is very alarming, and it is already going beyond the Exabyte limit. Hence, there is an obvious need for fast and accurate big data classification tools. Machine Learning (ML) based solutions are very useful and reliable data classification tools, however, they cannot effectively handle large-scale datasets. This paper therefore proposes two intelligent instance selection techniques for optimizing the training and classification speed of ML algorithms, with a specific focus on Support Vector Machine (SVM). Furthermore, this paper considers two different approaches to instance selection namely: filter-based and wrapper-based. Different sets of experiments are performed on 20 small-scale datasets and 10 large- or medium-scale datasets. The results show that the proposed techniques improved SVM training speed in 100% (30 out of 30) of the datasets used for evaluation and simultaneously improved SVM predictive accuracy in some cases. Furthermore, statistical analysis test is carried out and the results reveal that the training speed of the proposed techniques is statistically significantly faster than the training speed of standard SVM and some other existing instance selection techniques. In real life application, such as video surveillance and intrusion detection systems, that require a classifier to be trained quickly for speedy classification of new target concepts, the filter-based techniques provide the best solutions; while the wrapper-based techniques are better suited for applications such as email filters, that are very sensitive to slight changes in predictive accuracy.

Read the paper · More papers on PaperTik