Projected Clustering Particle Swarm Optimization and Classification

Satish Gajawada, Durga Toshniwal · 2012

Supervised learning algorithms are trained with labeled data only. But labeling the data can be costly and hence the amount of labeled data available may be limited. Training the classifiers with limited amount of labeled data can lead to low classification accuracy. Hence pre-processing the data is required for getting better classification accuracy. Full dimensional clustering has been used in literature as pre- processing step to classification methods. But in high dimensional data different clusters may exist in different subspaces of the dataset. Projected Clustering Particle Swarm Optimization (PCPSO) finds optimal centers of subspace clusters by optimizing a subspace cluster validation index. In this paper we use PCPSO method to find subspace clusters present in the dataset. The subspace clusters found and limited amount of available labeled data are used to label the large amount of unlabelled data that is present in the dataset. Various classification methods are then applied on the data pre-processed by using PCPSO. In this paper we propose PCPSO-Classification method. Various new classification methods like PCPSO-Naive bayes, PCPSO-Multi layer perceptron and PCPSO-Decision table can be obtained by using different classification methods like Naive bayes, Multi layer perceptron and Decision table respectively in the classification stage of proposed PCPSO-Classification method. When the dataset contains subspace clusters and labeling the data is costly due to which available labeled data is limited then the structure of data may be used along with available limited labeled data to label the large amount of unlabeled data. After pre-processing the data the amount of labeled data is not limited. We applied PCPSO-Naive bayes, PCPSO-Multi layer perceptron and PCPSO-Decision table methods on synthetic datasets and found classification accuracy improved significantly compared to using Naive bayes, Multi layer perceptron and Decision table for classification with limited available labeled data for training classifiers. The subspace clusters found by PCPSO can be used for different types of pre-processing for solving different problems before applying classification methods on datasets. In this paper we considered the problem of limited labeled data and using PCPSO to find subspace clusters which are used for labeling large amount of unlabeled data with the help of available limited labeled data.

Read the paper · More papers on PaperTik