Study on Algorithms for Local Outlier Detection

Xue An · Chinese Journal of Computers · 2007

Outlier detection has attracted much attention recently. There are two kinds of outliers: global outliers and local outliers. In many scenarios, the detection of local outliers is more valuable than that of global outliers. To mine local outliers, it is more meaningful to assign to each object a degree of being an outlier. Some existing representative algorithms currently used for solving this problem are compared in detail, and their disadvantages are pointed out such as poor efficiency and the detection accuracy depending on the parameters given by the user. In general, the attributes of each data object can be categorized as the inherent attributes and the context attributes, the inherent attributes characterize the data object while the context attributes embody the relationship between this data object and the neighbor data objects. The context attributes is not intrinsic to the data object. In order to overcome those disadvantages mentioned above, this paper proposes to use the context attributes to determine the object neighborhood and use the inherent attributes to compute the outlier score. For spatial data, the attributes comprise the non-spatial dimensions and the spatial dimensions. The spatial attributes provide a location index to the data object. The neighborhood in the Euclidean space plays a very important role in the analysis of spatial data. The spatial attributes are used to determine spatial neighborhood and the non-spatial dimensions are used to compute the spatial outlier score. This paper also proposes a novel measure, spatial local outlier factor (SLOF), which captures the local behavior of datum in its spatial neighborhood. The experimental results show that proposed SLOF algorithm outperforms the other existing algorithms in detection accuracy, user dependency, scalability and efficiency.

Read the paper · More papers on PaperTik