Exploring Multi-Label Data Augmentation for LLM Fine-Tuning and Inference in Requirements Engineering: A Study with Domain Expert Evaluation

Hanyue Liu, Marina Bueno García, Nikolaos Korkakakis · 2024

The application of Large Language Models (LLMs) has led to advancements in requirements engineering by providing automated solutions for various tasks. However, these models face challenges when processing complex and confidential multi-label data. This study investigates the impact of data augmentation on LLM performance during fine-tuning and inference in requirements engineering. A novel data augmentation technique is introduced specifically designed for multi-label technical data. The effectiveness of this approach is assessed through controlled experiments involving three fine-tuned LLMs, with evaluations conducted using clearly defined metrics. Three state-of-the-art LLMs were used to assess the performance of these models, and five domain experts validated the results. The findings show that: 1) correct implementation of the proposed data augmentation technique can improve LLM performance; however, incorrect implementation may have adverse effects; 2) the quality of training datasets limits the potential performance of fine-tuned LLMs; and 3) with a well-defined evaluation framework, LLMs can serve as effective judges in requirements engineering tasks, even without extensive domain-specific expertise. This research provides insights for developing more robust and efficient LLM-based frameworks in specific contexts.

Read the paper · More papers on PaperTik