An Outlier Mining Algorithm in High-Dimension Based on Single-Parameter-k Local Density

Weili Huang, Di Wu, Jiadong Ren · 2009

As one of the most important problems in data mining, many studies have been done on mining outliers. However, mining outliers in high-dimension has not been well addressed. In this paper, the concepts of reference radius and local deviation index are defined. A novel algorithm OMHKLD based on single-parameter-k local density in high-dimension for mining outliers is proposed. According to a new clustering algorithm KLDCA based on single-parameter-k local density, the data set is divided into outliers and cluster points. The cluster points are eliminated directly. The outlier candidate set is obtained. Moreover, take advantage of the idea of LOF, our algorithm indicates the degree of the objects in outlier candidate set with the local deviation index. The optimal outlier set can be gained. The experimental results and analysis show that the performance of OMHKLD is better than DBSCAN and LOF in improving the clustering quality and reducing memory usage and time cost.

Read the paper · More papers on PaperTik