Domain-specific data gathering and exploitation

Nabil Moncef Boukhatem, Davide Buscaldi, Leo Liberti · HAL (Le Centre pour la Communication Scientifique Directe) · 2024

This paper addresses challenges in domain-specific data acquisition for NLP, particularly where real-world data is limited due to privacy or availability constraints. Two primary approaches are explored: synthetic data generation using models like GANs and VAEs, and advanced data cleaning techniques to maximize existing datasets. Additionally, the role of large language models (LLMs) and multimodal models in automating data tagging, filtering, and preprocessing is discussed. Case studies in healthcare, finance, and cybersecurity demonstrate the effectiveness of these methods. The paper concludes by highlighting future directions, including scalability, privacy concerns, and the integration of LLMs in domain-specific applications.

Read the paper · More papers on PaperTik