FILTA: Better View Discovery from Collections of Clusterings via Filtering

Yaguo Lei, Xuan Phi Nguyen, Jason W. Chan, James A Bailey · 2014

Abstract. Meta-clustering is a popular approach to find multiple clus-terings in the datasest, which takes a large number of base clusterings as input for further user navigation and refinement. However, the effec-tiveness of meta-clustering is highly dependent on the distribution of the base clusterings and open challenges exist with regard to its stability and noise tolerance. In this paper we propose a simple and effective fil-tering algorithm (FILTA) that can be flexibly used in conjunction with any meta-clustering method. Given a (raw) set of base clusterings, FILTA employs information theoretic criteria to remove those having poor qual-ity or high redundancy. Then this filtered set of clusterings is highly suitable for further exploration, particularly the use of visualization for determining the dominant views in the dataset. We evaluate FILTA on both synthetic and real world datasets, and see how its use can enhance view discovery for complex scenarios.

Read the paper · More papers on PaperTik