The Use of Consensus Clustering in Geodemographics
James A. Cheshire, Muhammad Abdullah Adnan · UCL Discovery (University College London) · 2011
Geodemographic classifications require clustering algorithms to partition the records of large multidimensional datasets into groups sharing similar characteristics. Many clustering algorithms have been developed but few have been as widely implemented as the traditional methods such as K-means or Ward's hierarchical clustering (Jain, 2010). No two methods create the same result, and multiple iterations of the same method may produce different clusters; it is left to the user to subjectively decide the best outcome. In addition most methods require an a priori impression of the number of groups in the data. This abstract outlines a new approach, known as consensus clustering, that utilises familiar clustering methods to produce more consistent results. The method offsets the weaknesses of one type of clustering with the strengths of another by establishing the consistent average outcome from multiple algorithms (Simpson et al. 2010). Consensus clustering has an additional advantage in that it provides a number of metrics that inform the researcher about the inherent groups within the data, and the robustness of the final cluster outcome. Still in its early stages of development, and largely applied in the fields of genetics and bioinformatics, the method has some performance issues when using large datasets but we are confident these can be overcome.