Augmentation in a Binary Text Classification Task
Bohdan M. Pavlyshenko, Mykola Stasiuk · 2023
- The purpose of this paper is to investigate the effect of different kinds of augmentation on the binary text classification performed by different transformer models. Augmentations were performed in three ways: synonym augmentation, contextual word embeddings, and combined. For classification, BERT, ALBERT, DistilBERT, and RoBERTa transformer models were used. It has been found that when using context word embeddings augmentation every model started to overfit during the second training epoch. In the case of combined contextual word embeddings and synonym augmentations utilization, the overfitting issue was overcome, and models exhibited overall good performance. The best performance, however, was obtained when synonym augmentation was used: the overfitting issue was also avoided, and the models' effectiveness was the highest among all experiments.