Modelling a stable classifier for handling large scale data with noise and imbalance

S Akila, U. Srinivasulu Reddy · 2017

Classifier performance is often impaired by the presence of anomalies like noisy and borderlines samples, and due to the inherent imbalance in data. This is due to the fact that classifier models are usually constructed on the basis of ideal data conditions which is often not the case. In reality, these anomalies occur at varying intensities and in most cases they are an integral part of the problem domain. This requires that the classifier models be fine-tuned to accommodate such anomalies thereby resulting in data dependent models. This work analyses the effectiveness of various classifier models in handling noisy, borderline and imbalanced data. This dictates that, the right set of metrics must first be identified, as most of the usual metrics are not affected by such anomalies, though it affects the reliability, robustness and practical efficacy of such classifiers. To ensure the scalability of the resulting models, classifiers were implemented using Spark. A characterized examination of the results elucidates the effective prediction zones of each model, facilitating the identification of stable classifier models. It is found that a single model is inadequate in real time scenarios, due to the complex interplay among the various anomalies. This work is concluded with a modelling a heterogeneous cost based ensemble model for a domain based prediction model.

Read the paper · More papers on PaperTik