Different Data Enhancement Methods applied on Imbalanced Data

Yongkai Deng, Chenyu Fan, Sutong Jin, Hongxi Zhang · 2022

Imbalanced data distribution caused over-centered performance when doing humorousness rating. The majority of data focuses on a specific part instead of varies at an appropriate range. This paper aims at implementing an ideal combination of different methods to improve the final result of the deep learning model. We tried various methods to smooth the skewed distribution. The results show that the best collaboration of methods is kernel density estimation, synonym substituting, sentence paraphrasing, under-sampling to develop the model performance. After testing more than five times, it could be seen that the model performs a little better than the model designed previously, no matter from which aspects. Although the validation loss and training loss was controlled at an acceptable error, we still encountered some dreadful challenges like missing predictions of the actual data, the severely skewed distribution of data predictions.

Read the paper · More papers on PaperTik