A Study of Anomaly Detection Techniques

Sanketh Harnoorkar · International Journal for Research in Applied Science and Engineering Technology · 2020

Anomaly detection refers to identification of data items, points or events that are rare, differ significantly from other data items, points or events or that have unexpected behaviour.These rare items are called anomalies, outliers, exceptions, defects or contaminants.Anomaly and outliers are the 2 commonly used words. Two assumptions have to hold off for the anomaly detection to be effective -number of normal patterns should be higher than the no of anomalies and anomalies should be distinguishable from normal patterns. Anomalies or outliers were detected as a part of cleaning the data. However, it was later found in 2000 that detection of anomalies can help in solving real world problems. Anomaly detection can help solving problems such as intrusion finding, fraud detection in credit card transactions, system health monitoring and industrial damage detection. The anomalies are classified into point, contextual and collective anomalies. If it's found that only a single data point differs significantly from other data points in terms of attributes, then it is called a point anomaly. The main aim of anomaly detection is to detect cases which are not usually found within data that is of a similar kind.. Anomalies in the data can occur for different reasons.. Data analysts find these anomalies intriguing and interesting. . The bottom line is increase up of the time and the reduction of any downtime. The paper makes an attempt to analyse the various anomaly detection techniques by bringing out a survey of the existing research which exists in this domain.I. ANOMALY DETECTION ON TIME SERIES DATA Data from the time series are data which are collected at different times.This is against cross-sectional data which at a single point in time observes individuals , companies, etc. Due to the fact that data points are collected in time series at adjacent time periods, correlation between observations is potential.This is one of the features that sets time series data apart from cross-sectional.Time series data is often seen in a variety of domains in today's world.They are seen in the field of Economics, Epidemiology, Social Sciences, ,Medicine, Physical sciences.[1] explores the isolation forest method to detect anomalies and measure how severe they are.It emphasises how the present anomaly detection strategies have their shortcomings in which many of them set an arbitrary threshold for detecting anomalies and the presence or absence or anomalies is determined based on how close the present feature is close to the threshold.The paper talks about the shortcomings of this approach and talks of a better alternative to this in the form of random forests.This is useful whenever there are many anomaly points and there is confusion on where to start the search from.The paper also says that the research is backed by experimental evidence.It deals with the main problem on finding out the threshold of anomaly.It proposes a k means clustering to decide on an adaptive algorithm to divide between outliers and normal values.It initially proposes a normal k means algorithm on two clusters and then proposes some modifications and enhancements to it and gives an anomaly score to each of the outlier detected.The methodology includes using an adaptive algorithm based on feature points, pattern feature extraction and mapping to a feature space.After this the isolation forest part comes in where with different tree construction algorithms a random forest is constructed.After this, anomaly score is calculated and anomaly sets are clustered based on K-means.After this we calculate the K nearest points of each point.According to the paper, experiments prove the Improved iForest algorithm is better suited for use in Big hydrological time series data and high characteristics precision in detecting anomalies.. [2] This paper deals with seasonal time series data and proposes an unsupervised algorithm.This proposed a look back window and takes a subset and takes a median as expected value.The actual value of new data is considered and compared with the median, it also proposes a percentile window and also a 3sigma detector.They've also suggested a validation procedure.They've named it as MULDER.According to the results they've obtained results better than amazon and twitter's algorithm.They've proposed the algorithm mainly for streaming systems.They've also laid out the challenges presently with streaming services.Mainly the problem with not having test/validation sets, not having a set normal defined, difficulties in accurately naming anomalies mainly due to their volume, velocity and variety of data and also manual parameter tuning.[4] deals with anomaly detection on time series data using markov chain.It first details about feature extraction, which involves transforming a pressure data segment into markov chain, one step probability, amplitude information of time series and combined feature vector.The dataset is obtained from Supervisory Control And Data Acquisition (SCADA) system from Tinjin, China.In the paper, a new feature extraction method is introduced and validated with the dataset.

Read the paper · More papers on PaperTik