Contrast mining in large class imbalance data
Jingyuan Li · UTS ePRESS (University of Technology Sydney) · 2013
Class imbalance data, in which the classes are not equally represented and the minority classes include a much smaller number of examples than other classes, is pervasive and ubiquitous, particularly in applications such as fraud/intrusion detection, medical diagnosis/monitoring, and risk management.The conventional classifiers tend to be overwhelmed by the large classes while ignoring the smaller classes.Typically, many of the existing solutions to the class imbalance problem are proposed at the data level, and a few at the algorithmic level.However, the prior methods have more or less limitations in anomaly detection according to our extensive experiments.Therefore, the thesis targets contrast mining to solve the problem of anomaly detection in imbalanced data from three aspects: feature construction, an effective algorithm for mining contrast patterns, and selection of optimal rule combinations through analysing rule interactions.Feature construction is one of the most important steps in contrast pattern mining, and any other data mining processes as well.The majority of feature construction methods, such as Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA), Fourier Transformation, and Independent Component Analysis, usually generate new features by transforming the existing raw features into a new data space.Therefore, previous solutions have many limitations with respect to the objective of training highly accurate classifiers in class imbalance data sets.Incomprehensible features may be generated, based on the assumption that all the samples are independent, the feature set is unstable and sensitive to trivial change of the sample set, xii ABSTRACT it is difficult to integrate significant domain knowledge, and the classifiers built on the transformed feature set suffer from high False Positive Rate in the class imbalance data set.In order to train high performance models in the imbalance scenario, we propose a novel method, Personalised Domain Driven Feature Mining (PDDFM ), to generate important features by integrating domain knowledge effectively with a full consideration of the correlations among samples.A framework specially designed for PDDFM is introduced.A novel feature selection method, called Mutual Reduction, is proposed to minimise the noise from redundant features and maximize the contribution of "trivial" features whose gain ratio are low but contribute positively when cooperate with the others.The experimental evaluation reveals our feature mining approach outperforms state-of-the-art methods in anomaly detection.Contrast pattern mining has been studied intensively for its strong discriminative capability.However, state-of-the-art methods rarely consider the class imbalance problem, which has been proven to be a significant challenge in mining large scale data.The thesis introduces a novel pattern, i.e. converging pattern, which refers to the item sets whose supports contrast sharply from the minority class to the majority class.A novel algorithm, ConvergMiner, is also proposed to mine converging patterns efficiently.A light-weighted index T*-tree is built to speed up the search process, and output patterns instantly.A series of branch bound pruning strategies are further presented to greatly reduce the computational cost.Substantial experiments on large scale real-life online banking transactions for fraud detection show that the ConvergMiner greatly outperforms the existing cost-sensitive classification methods in terms of accuracy.In particular, it efficiently and effectively detects the frauds in large-scale imbalanced transaction sets.More importantly, the efficiency improves with the increase in data imbalance.After many converging patterns are generated, we propose an effective novel method to select the optimal pattern set.Rule-based anomaly and fraud detection systems often suffer from subxiii