A critical analysis of Sampling Techniques for imbalanced data classification: An application to Social Media

Vinitha Nagarajan · NORMA · 2017

Imbalanced training datasets appear in a number of real-life problems, such as anomaly detection, network monitoring and social media. While some classes will normally have a sizeable amount of records, other classes will be underrepresented. Constructing efficient classifiers for minority classes is a challenge which has been addressed in various ways, but basically grouped into undersampling the majority class(es) and oversampling the minority class(es), or a combination of both techniques. This thesis will focus on imbalanced classification techniques for social media data from Twitter. The classification task at hand is the identification of spam tweets. Social networks are a rich source of information, but also attracts many illegitimate users who spread spam tweets. A machine learning approach is presented whereby analytical models are learnt from highly skewed datasets to predict spam messages. A range of techniques to tackle class imbalance are analysed in detail by controlling the class imbalance ratio. It is indeed possible to identify techniques of superior performance according to the imbalance ratio they can cope with. It is shown that classification performance heavily depends on the imbalance degree of the dataset and their sampling techniques.

Read the paper · More papers on PaperTik