Large-Scale Instance Selection using Center of Principal Components
Chatchai Kasemtaweechok, Patnaree Pharkdepinyo, Patimakorn Doungmanee · 2021
One-hot encoding is one of the most popular data preprocessing methods used to convert categorical data into numerical data before applying a training classification model. Drawbacks of one-hot encoding are its high memory consumption and large storage requirements when the training set has many categorical columns. To overcome these drawbacks, we proposed the PC-CS method to reduce the training set size by only selecting principal components center as a representative instance of each disjoint partition. The proposed method was compared with two instance selection methods and the full training model as the baseline model. We used five datasets from the UCI dataset repository website. Evaluation of the classification performance involved four classifier algorithms: decision tree, naive Bayes, support vector machine, and logistic regression. The average classification performance of the proposed PC-CS method was approximately 2% higher than those of the other algorithms tested, while the reduction rate was about 5% higher.