Improving Clustering Performance by Using Feature Selection and Extraction Techniques
Khalil Shihab · Journal of Intelligent Systems · 2004
Clustering of data is of great interest in many data mining applications.In applying a clustering technique, however, many attributes or features can be irrelevant to the clusters produced.Therefore, reducing the number of dimensions has proven to be a valuable technique for improving the efficiency of a clustering algorithm, especially when the input data vectors contain a large number of features.Feature selection and extraction techniques aim at selecting a subset of the features that is relevant for a given problem.Usually all features do not generate a corresponding increase in performance of the clustering method.Some of these features may be noisy, meaningless, correlated, or irrelevant for the clustering task.In particular, the attributes that have similar data across the majority of components (data vectors) should be deleted because these attributes do not add any useful information for producing different clusters.In this work, we apply feature selection and feature extraction techniques as a pre-processor for our proposed conceptual clustering method to improve clustering performance and to reduce its computational complexity.The results presented from the application of these methods to computer workload characterization in particular indicate that the integration of feature selection and extraction methods with conceptual clustering has potential for producing meaningful categories.