GCA: An algorithm based on the gower similarity for clustering of categorical variables
Tiago Rodrigues Lopes dos Santos, Luis Enrique Zárate · 2012
The data clustering is a technique used to make groups of objects present similar characteristics from a database. These databases may contain different variable types (numeric, categorical, scalar, binary, etc..), but categorical variables such as become a challenge clustering because lack of natural ordering. With this lack there is a big deficiency of tools and algorithms for clustering databases with categorical variables. The present work propose a new clustering algorithm for categorical data called GCA (Gower Clustering Algorithm) based in combination of algorithm TaxMap and measure of similarity coefficient of Gower. The GCA algorithm was compared with two others algorithms (clope and FarthestFirst) and through a brief statistical analysis, GCA had a very significant performance to contribute with deficiency cited.