Binary-based similarity measures for categorical data and their application in Self- Organizing Maps
Fernando Correia Lourenço, Victor J. A. S. Lobo, Fernando Bação, Gestão de Informação · 2004
In exploratory data analysis of high dimensional data one Eof the main tasks is the formation of a simplified overview of data sets. Clustering and projection are among the examples of useful methods to achieve this task. However there are several types of data where the use of this measure is not adequate, such as the categorical data. In this paper we will review some of the most common binary-based similarity measures that can be applied to this type of data. These measures are evaluated empirically using the Self-Organizing Maps (SOM) algorithm. The SOM algorithm performs a non-linear mapping from a high-dimensional data space to a low-dimensional space, typically twodimensional, aiming to preserve the topological relations of the data. A well known data set of Animals from (Ritter and Kohonen 1989) is used. We tested two different approaches to compute the best matching unit (BMU). Working with binary data in the SOM algorithm poses issues concerning the update method for the neurons and the internal modeling of the data, which we have to be aware. The nonmetric properties of some measures must also have our attention. Some similarity measures produced maps that were very like to each other. Exploring these maps, we may find that the clustering obtained using the SOM provides different perspectives over the data.