Large-scale clustering using mathematical programming

Mario Gnägi, Philipp Baumann · 2017

Cluster analysis is a fundamental task in exploratory data analysis with a wide range of applications. Several clustering approaches based on mathematical programming have been proposed in the literature and were successfully used for small- and medium-scale data sets. However, mathematical programming-based clustering models are rarely used for large-scale data sets due to their extensive running time. In this paper, we propose a general scaling approach for existing mathematical programming-based clustering models that is based on the idea of replacing identical or nearly-identical objects by a small set of representatives. Our computational results indicate that the proposed scaling approach substantially reduces running time with a minor loss in clustering accuracy.

Read the paper · More papers on PaperTik