Systematical Randomness Assignment for the Level of Manipulation in Text Augmentation

Youhoo Cha, Young‐Hoon Lee · 2024

Text augmentation, a method for generating new texts by using noise, combinations, and other mixings to a scarce dataset, is an important skill in natural language processing (NLP). This allows for the introduction of diversity into the training process, resulting in more robust models. However, in the field of text augmentation, the level of manipulation can cause the following problems: When manipulation is 'low-level', it cannot guarantee diversity by generating data similar to the original but can lead to inefficient augmentation, while 'high-level' manipulation causes unreliable label issues and degrades model accuracy. Therefore, in this paper, a text augmentation technique is proposed by systematically assigning randomness to solve the “level of manipulation” problem. Additionally, we generate an advanced sentence embedding that can assign robust pseudo-labels at a high manipulation level. That is, advanced sentence embeddings capable of assigning reliable pseudo-labels are generated by extracting information from the original data, namely sentence embeddings, document embeddings, and eX-plainable Artificial Intelligence(XAI) information. We verify the effectiveness of the proposed methodology through sentiment classification accuracy comparisons with existing text augmen-tation approaches, and show that the proposed methodology achieves high sentiment classification accuracy improvements on most experimental datasets. Index Terms-text augmentation; advanced sentence embed-ding; reliable pseudo-labels; the level of manipulation

Read the paper · More papers on PaperTik