Applying Average Density to Example Dependent Costs SVM based on Data Distribution
Xin Yu Jin, Yujian Li, Yi‐Hua Zhou, Zhi Cai · Journal of Computers · 2013
Standard Support Vector Machines (SVM) often performs poorly on imbalanced datasets, because it could not get a high accuracy of prediction on the minority class of data as well as another class. We proposed a new example dependent costs SVM method, from which it can get more sensitive hyperplane by selecting penalty for every sample according to its corresponding distribution. Firstly, this paper discusses how to create an Example Dependent Costs SVM based on Data Distribution (DDEDC-SVM), and then we proposes a direct method to determine the parameters, i.e., “Average Density”, in order to reduce the time for their selection via traditional cross validation. Experimental results show that this method can improve the performance of SVM on imbalanced datasets efficiently and effectively.