FOFO: Fused Oversampling Framework by addressing Outliers

Arghasree Banerjee, Kushankur Ghosh, Sankhadeep Chatterjee, Diptaraj Sen · 2021 International Conference on Emerging Smart Computing and Informatics (ESCI) · 2021

The challenges regarding data distribution are proved to be fatal for any intelligent model. The imbalanced distribution among data is acknowledged to be a popular problem in real-life datasets. The disproportion in data distribution is often associated with the presence of outliers in the dataset. However, Machine Learning models have not been tested properly with synthetic sampling in presence of irregular data. In our paper, we propose a data pre-processing framework that follows a selective approach. A 2-step oversampling is done on the outliers and the entire dataset is done by effectively utilizing the Synthetic Minority Oversampling Technique (SMOTE) to mitigate data biases and the effects of outliers. Our model was tested with 6 robust machine learning models and on 3 real-life datasets. We further compared our framework with existing oversampling techniques. It is found that the proposed approach is has improved the performance of each of the classifiers.

Read the paper · More papers on PaperTik