Clustering of large data based on the relational analysis
Said Chah Slaoui, Yasmine Lamari · 2015
This paper presents a fast heuristic which finds clusters by partitioning categorical large data sets according to the Relational Analysis, whereby the cluster analysis is modeled as a linear integer program with n2attributes (n is the number of observations) and solved by the optimization under constraints of the Condorcet criterion. Without neither a sampling method nor the fixing of input parameters and while using a natural cluster structure, Transitive heuristic needs a small amount of memory and a short time to provide good quality partition. Experimental results on real and synthetic data sets are presented in order to show that clusters, formed using this technique, are intensive and accurate.