Unsupervised clustering of multidimensional distributions using earth mover distance
David Applegate, Tamraparni Dasu, Shankar Muthu Krishnan, Simon Urbanek · 2011
Multidimensional distributions are often used in data mining to describe and summarize different features of large datasets. It is natural to look for distinct classes in such datasets by clustering the data. A common approach entails the use of methods like k-means clustering. However, the k-means method inherently relies on the Euclidean metric in the embedded space and does not account for additional topology underlying the distribution.