Parallel Algorithm for Mining Outliers in Large Database

Edward Hung, Dwl Cheung · 1999

Data mining is a new, important and fast growing database application. Outlier (exception) detection is one kind of data mining, which can be applied in multiple areas like monitoring of credit card fraud and criminal activities in electronic commerce. With the ever-increasing size and attributes (dimensions) of database, previously proposed detection methods for two dimensions are no longer applicable. The time complexity of the Nested-Loop (NL) algorithm is linear to the dimensionality but quadratic to the dataset size, inducing an unacceptable cost for large dataset. A more efficient version (ENL) and its parallel version (PNL) are introduced. ENL reduces the cost to half while the cost in PNL is quadratic to the reciprocal of the number of processors. There is also a performance comparison between ENL and PNL using Bulk Synchronization Parallel (BSP) model. Keywords: Data Mining, Outlier Detection, Parallel Algorithm 1 Introduction Data mining or knowledge discovery tasks can be ...

Read the paper · More papers on PaperTik