Synthetic Textual Data Generation: A Few-Shot Learning-Based Approach for DPPs with Novel Metrics
A. M. Esfar-E-Alam, Amir Taherkordi · 2025
The European Union aims to adopt Digital Product Passports (DPPs) to advance circular economy principles by facilitating the reuse and tracking of materials. However, the scarcity of available DPP data presents a significant challenge. Specifically, limited standardization, data-sharing constraints and confidentiality concerns restrict the accessibility and uniformity of DPP information, hindering effective software model development. This research explores open-ended text generation, focusing on synthetic DPPs, using Large Language Models (LLMs) through few-shot learning to accurately replicate DPP attributes. We evaluate models such as GPT-2, GPT-4oMini, LLaMA 3.2 and Gemma 2B using BLEURT, COMET and an innovative Hybrid Composite Analysis (HCA) score. Our findings indicate that few-shot learning is an optimal method for generating realistic synthetic data, with GPT-4oMini delivering the best performance. While BLEURT and COMET provide useful benchmarks, they often diverge from human preferences in evaluating open-ended, structured text data generation. To address this gap, we propose a novel evaluation approach tailored to structured yet flexible outputs like DPPs, demonstrating its effectiveness in aligning with human judgment and aiding sustainability initiatives.