How to Use the Fractal Dimension to Find Correlations between Attributes
Elaine Parros Machado, Agma J. M. Traina, Christos Faloutsos · 2002
One of the challenges when dealing with multimedia information, which usually is massive and composed of multidimensional data, is how to index, cluster and retain only the relevant information among the items of a large database. In order to reduce the information to the meaningful data, techniques such as attribute selection have been used. In this paper we present a fast, scalable algorithm to quickly select the most important attributes (dimensions) for a given set of n-dimensional vectors, determining what attributes are correlated to the others and how to group them. The algorithm takes advantage of the ‘fractal’ dimension of a data set as a good approximation of its intrinsic dimension and, based on it, indicates what attributes are the most important to keep the meaning of the data. We applied our method on real and synthetic data sets, producing fast and promising results.