Notice of Violation of IEEE Publication Principles Improving heterogeneous data clustering by using metadata and compression algorithms
Alexandra Suzana Cernian, Dorin Cârstoiu, Valentin Sgârciu · 9th RoEduNet IEEE International Conference · 2010
Nowadays, we have to deal with a large quantity of unstructured, heterogeneous data, produced by an increasing number of sources. Clustering heterogeneous data is essential to getting structured information in response to user queries. In this paper, we assess the results of a new clustering technique - clustering by compression - when applied to metadata associated with heterogeneous sets of data. The clustering by compression procedure is based on a parameter-free, universal, similarity distance, the normalized compression distance or NCD, computed from the lengths of compressed data files (singly and in pair-wise concatenation). Experimental results show that using metadata could improve the average clustering performances with about 20% over clustering the same sample data set without using metadata.