An LLM-Based Framework for Synthetic Data Generation
Mandeep Goyal, Qusay H. Mahmoud · 2025
The demand for high-quality datasets is rapidly increasing across sectors such as healthcare, finance, and cybersecurity, yet challenges like data scarcity and privacy concerns persist. To address this, we introduce a framework for synthetic data generation that empowers users to create realistic datasets while maintaining privacy. The framework leverages fine-tuned Large Language Models (LLMs) and differential privacy techniques, including IBM's diffprivlib, to generate synthetic data that replicates real-world patterns without exposing sensitive information. A proof-of-concept platform has been constructed to facilitate seamless data generation and augmentation, making it particularly useful in scenarios where original datasets are inaccessible, scarce, or privacy-restricted. The platform supports the creation of datasets across five key categories, employing advanced methods to preserve data integrity while ensuring compliance with stringent privacy standards. By combining cutting-edge AI technologies with robust privacy-preserving techniques, this framework offers a practical solution for researchers and professionals seeking reliable synthetic data to drive innovation in data-sensitive fields.