Text Generation for Imbalanced Text Classification
Suphamongkol Akkaradamrongrat, Pornpimon Kachamas, Sukree Sinthupinyo · 2019
The problem of imbalanced data can be frequently found in the real-world data. It leads to the bias of classification models, that is, the models predict most samples as major classes which are often the negative class. In this research, text generation techniques were used to generate synthetic minority class samples to make the text dataset balanced. Two text generation methods: the text generation using Markov Chains and the text generation using Long Short-term Memory (LSTM) networks were applied and compared in the term of ability to improve the performance of imbalanced text classification. Our experimental study is based on LSTM networks classifier. Traditional over-sampling technique was also used as baseline. The study investigated our Thai-language advertisement text dataset from Facebook. According to the increase of recall value, applying of these techniques showed the improvement of an ability to create model predicting more positive samples, which are minority samples. It can be found that the Markov Chains technique outperformed traditional over-sampling and text generation using LSTM in majority of the models.