EnvBERT: Multi-Label Text Classification for Imbalanced, Noisy Environmental News Data

Dohyung Kim, Jahwan Koo, Ung Mo Kim · 2021

Imbalanced and noisy classification problems pose a challenge for predictive modeling as most of the machine learning algorithms used for classification were designed around the assumption of an equal number of non-noisy examples for each class. Models with these problems cause classification errors. We propose a multi-label text classification model based on BERT, EnvBERT, which includes multi-label features in text classification and has good predictive performance for imbalanced, noisy environmental news data. EnvBERT is based on the KoBERT model pre-trained with Korean text data. We used the data oversampling technique to resolve the imbalanced characteristics of multi-label data and fine-tuned while setting a global threshold for label prediction. As a result, we show that EnvBERT improves classification performance by more than 80% on the imbalanced and noisy environmental news data.

Read the paper · More papers on PaperTik