NES: Neural Embedding Squared

Lucas Zanco Ladeira, Frances Albert Santos, Leandro Aparecido Villas · Journal of Internet Services and Applications · 2025

In the fields of natural language processing (NLP) and machine learning, the quality and quantity of training data play a pivotal role in model performance. Textual data augmentation, a technique that artificially enhances the size of the training dataset by generating diverse yet semantically equivalent samples, has emerged as a crucial tool for overcoming data scarcity and improving the robustness of NLP models. However, the available solutions that achieve state-of-the-art performance require considerable computing power. This occurs because they use resource hungry machine learning models for each synthetic sentence generated. This paper introduces an approach to textual data augmentation, leveraging semantic representations to produce augmented data that not only expands the dataset but also understands the distribution of data points spatially. The approach requires less computing power by exploiting a fast prediction and spatial exploration in the embedding representation. In our experiments, it was able to double model performance while fixing class unbalance.

Read the paper · More papers on PaperTik