Efficient Acronymization of Sensitive Data Using Generative Models Fine-Tuned on LLM-Augmented Data
Kristián Sopkovič, Eva Kupcová, Daniel Hládek, Matúš Pleva · 2025
Data privacy is crucial today, especially with regulations such as GDPR. Data anonymization is key, but common methods often reduce data value. This paper explores the acronymization process, which replaces sensitive data with abbreviations, as a way to balance protection and data usability. We propose a new approach: using large language models (LLMs) to create training data for a T5 model, which we then fine-tune for acronymizing sensitive data. The results show that this method, especially when using LLMs like Gemma-9B-IT to augment the data, achieves promising results and outperforms existing Named Entity Recognition (NER) models in the specific task of acronymization. This offers a more efficient and scalable solution for anonymizing text data, contributing to both privacy protection and preserving the utility of data for analysis and research.