Adaptive Data Augmentation Techniques for Software Requirements Classification Using Deep Learning

Osamah AlDhafer, Hussain Alsadiq, Mohammed Alshaboti · IEEE Access · 2026

The scarcity and imbalance of labelled training data in real-world requirement engineering datasets constrain the effectiveness of automating software requirement classification using deep learning. This paper presents a suite of adaptive data augmentation techniques specifically designed to enhance deep learning performance for software requirement classification tasks. This study proposes novel augmentation strategies that use the class distribution and length of requirements to dynamically determine the number and proportion of synthetic samples to be generated. The proposed methods include weighted augmentation for one-label and multilabel scenarios, as well as semantic augmentation using domain-specific dictionaries and transformer-based language models. The techniques are integrated into deep learning pipeline based on Bidirectional Gated Recurrent Unit (BiGRU) and evaluated on two datasets PROMISE and EHR across binary, multiscale, and multilabel classification. The experimental results showed that the proposed method significantly improved classification performance, particularly for minorities and underrepresented classes where traditional models struggle. Specifically, adaptive augmentation proposed an improved F1-score of up to 9.3% for underrepresented classes and increased the macro-average F1 by 6.7% on PROMISE-10-class tasks. Transformer-based augmentation outperformed all baselines, achieving an F1-score of 84.1% in the PROMISE multilabel setting compared to 77.5% without augmentation. Similarly, dictionary-based augmentation yielded a 5.2% gain in the EHR multilabel classification. The proposed approach requires minimal manual intervention, which makes it practical and scalable for real-world applications.

Read the paper · More papers on PaperTik