Learning Discriminative Representations for Big Data Clustering Using Similarity-Based Dimensionality Reduction
Nikolaos Passalis, Anastasios Tefas · 2018
Discriminative Clustering techniques simultaneously perform clustering and learn a representation that encourages the separability of the clusters. However, methods with high discriminative power tend to decrease clustering accuracy, since the cluster assignments are usually noisy. In this paper, a similarity-based dimensionality reduction method, that allows for learning regularized clustering-oriented representations and is able to efficiently scale to large datasets, is proposed. We avoid the pitfalls of highly discriminative methods, such as the Linear Discriminant Analysis (LDA), by maintaining a small similarity between the inter-cluster samples and a small dissimilarity between the intra-cluster samples instead of collapsing the intra-cluster samples and pushing the clusters as far apart as possible. Three datasets are used to demonstrate the ability of the proposed method to learn robust representations that improve the quality of the obtained clustering solutions over other clustering techniques.