Gender Swapping as a Data Augmentation Technique: Developing Gender-Balanced Datasets for Ukrainian Language Processing

Olha Nahurna, Mariana Romanyshyn · 2025

This paper presents a pipeline for generating gender-balanced datasets through sentencelevel gender swapping, addressing the genderimbalance issue in Ukrainian texts.We select sentences with gender-marked entities, focusing on job titles, generate their inverted alternatives using LLMs and human-in-the-loop, and fine-tune Aya-101 on the resulting dataset for the task of gender swapping.Additionally, we train a Named Entity Recognition (NER) model on gender-balanced data, demonstrating its improved ability to recognize gendered entities.The findings unveil the potential of genderbalanced datasets to enhance model robustness and support more fair language processing.Finally, we make a gender-swapped version of NER-UK 2.0 and the fine-tuned Aya-101 model available for download and further research.

Read the paper · More papers on PaperTik