Easy Data Augmentation for Handling Imbalanced Data in Fake News Detection
Rilo Chandra Pradana, Setiawan Joddy, Abba Suganda Girsang · 2023
Detecting fake news in the digital era is challenging due to the proliferation of misinformation. One of the crucial is-sues in this domain is the inherent class imbalance, where genuine news articles significantly outnumber fake ones. This imbalance severely hampers the performance of machine and deep learning models in accurately identifying fake news. Consequently, there is a compelling need to address this problem effectively. In this study, we delve into fake news detection and tackle the critical issue of imbalanced data. We investigate the application of Easy Data Augmentation (EDA) techniques, including back-translation, random insertion, random deletion, and random swap to mitigate the adverse effects of imbalanced data. This study focuses on employing these techniques in conjunction with a deep learning framework, specifically a Bidirectional Long Short-Term Memory (BiLSTM) architecture. The results of the EDA techniques will be systematically compared to see their effectiveness and their impacts on model performance. This study reveals that various EDA techniques, when coupled with a BiLSTM architecture, yield significant improvements in fake news detection. Among the experiments, it shows that Random Insertion, with an impressive accuracy rate of 81.68%, a precision score of 89.38%, and an F1-Score of 87.77% emerges as the most promising technique. The study also highlights the exceptional potential of Back-translation stands out with an 87.16% recall performance.