Sampling for Large-Scale Clustering
Olivier Bachem · Repository for Publications and Research Data (ETH Zurich) · 2018
Scaling clustering algorithms to massive data sets is a non-trivial task.Existing approaches often rely on sampling data points uniformly at random from the full data set.However, such approaches may fail as real-world data is often imbalanced and uniform sampling becomes inadequate.A natural remedy to this problem is to replace uniform sampling by nonuniform sampling.In this dissertation, we consider three different research directions where sampling is useful to scale clustering to massive data sets.I would like to express my sincere gratitude to all the people that have supported me over the years in the endeavour culminating in this dissertation.First and foremost, Prof. Dr.