Comparative analysis of density based outlier detection techniques on breast cancer data using hadoop and map reduce

Sourajit Behera, Rinkle Rani · 2016

Advancement of technology, has furnished several terabytes of data for companies which can be effectively summed under Data Mining. Finding useful pieces of information from such huge data has been the need of the hour. A term called Anomaly Detection [8] is used in the pretext to refer to data objects which do not confer to a notion of normal data objects. There are various density based clustering algorithms[10] used to categorize data objects as normal or anomalous by finding clusters within the data set. LOF[18] finds the anomalous data objects by finding local density of data objects with respect to local density of its neighbors. DBSCAN finds anomalous data objects by finding data objects surrounded by data objects (density) which are far away from the concerned data object. OPTICS an extension of DBSCAN finds clusters of arbitrary sizes. DENCLUE uses a set of density distribution functions. This paper shows the comparison of the density based algorithms i.e. LOF, OPTICS, DBSCAN, DENCLUE based upon parameters such as time taken on single cluster hadoop, noise accuracy detection level, number of anomalous instances detected on high dimensional data, handle varied density, input parameters and complexity etc.

Read the paper · More papers on PaperTik