Uncertain data mining: A review of optimization methods for UK-means

Swati Aggarwal, Nitika Agarwal, Monal Jain · International Conference on Computing for Sustainable Global Development · 2016

Real world data generally deals with uncertainty. The complexity of data uncertainty poses many challenges. The most widely used k-means clustering algorithm is used for cluster analysis. However, it doesn't deal with uncertain data. The UK-means algorithm, a modification of k-means handles uncertain objects whose locations are represented by probability density functions (pdfs). It computes expected distances(EDs) between objects' pdf and cluster representatives using costly numerical integrations. This paper provides a review about the current state-of-the-art pruning techniques in order to improve the efficiency and countervail the computational complexity of UK-means based upon extensive research performed during the last decade. These are Min-Max bounding box (BB) based technique which uses two ways to calculate bounds: metric and trigonometric. Then, Voronoi based techniques to reduce ED calculations and indexing using R-Tree to reduce pruning overheads are discussed. Finally, hybrid techniques are observed to be feasible for real world employment.

Read the paper · More papers on PaperTik