Multi-type Relational Clustering Approaches: Current State-of-the-Art and New Directions

Tao Li, Sarabjot Singh Anand · Warwick Research Archive Portal (University of Warwick) · 2008

The proliferation of multi-type relational datasets in a number of important real-world applications and the limitations resulting from the transformation of such datasets to fit propositional data mining approaches have led to the emergence of the discipline of multi-type relational data mining. Clustering is an important unsupervised learning task aimed at discovering structure inherent in data. In this paper, we survey the state-of-the-art in the field of relational clustering, providing a taxonomy of approaches and review some of the most representative algorithms within each category. We also present DIVA, our general framework for multi-type relational clustering, which combines the use of Representative Objects with multi-phase clustering in a bid to provide flexibility, efficiency and effectiveness in clustering relational datasets. Theoretical analysis and experimental results prove that our approach is more effective and efficient than a number of other algorithms proposed in literature.

Read the paper · More papers on PaperTik