An improved incremental training approach for large scaled dataset based on support vector machine
Jingcai Guo · 2016
The Support Vector Machine(SVM) is well known in machine learning and artificial intelligence for its high performance in data classification, regression and forecasting. Usually for large scaled dataset, an incremental training algorithm is applied for tuning or balancing the training cost and the accuracy in SVM applications. This paper presents an improved incremental training approach for large scaled dataset on SVM. We focus on data's own distribution information to unfold our research, we proposed a self adaptive clustering method to extract the area and density information of data, a border detection technologies and uncertainty strategy is applied to maintain the border and some potential samples. Our proposed method can greatly reduce the training error for incremental training on SVM, especially for some uneven distribution dataset. We can greatly tuning or balancing the training cost and the accuracy of algorithms to achieve a better performance.