On the Effectiveness of Zero-Shot and Few-Shot Pretrained Language Models for Software Requirement Classification

Md Shafikuzzaman, Md Rakibul Islam, Shuaib Zaman, Andrew Ma, Anwarul Islam Sifat · IEEE Access · 2025

Accurate classification of software requirements, particularly distinguishing between fun- ctional and non-functional requirements, is essential for enabling traceability, guiding system design, and ensuring project success. Traditional supervised machine learning methods have shown promise in this domain but suffer from limited generalizability, high labeling costs, and overfitting. Recent advances in pretrained language models (PLMs) offer new opportunities for effective requirement classification through zero-shot (ZS) and few-shot (FS) learning. This study investigates the effectiveness of ZS and FS PLMs, including prompt-based generative large language models (e.g.,GPT-4oandLlama3), in classifying functional vs. non-functional requirements and non-functional subtypes, such as usability, security, operational, and performance requirements. We conduct a comprehensive empirical evaluation across five classification tasks using two widely adopted datasets: PROMISE and SecReq. The study compares six PLMs, including two ZS, two FS transformer models, and two generative Large Language Models under both ZS and FS settings. We benchmark these models against state-of-the-art and domain-specific baselines, and complement the quantitative analysis with a detailed qualitative error analysis. FS models, such asAll-MPNET,All-DistilRoBERTa, andGPT-4oin the FS setting, consistently outperform ZS models across all tasks, achieving superior performance in multi-class NFR and security requirement classification. WhileNoRBERTremains superior in binary FR/NFR classification, our FS models outperform or match its performance in other scenarios. ZS models, such asGPT-4o, also demonstrate strong results without requiring training data, highlighting their practical utility. Error analysis reveals that annotation issues and lack of context are major sources of misclassification. Our findings establish that ZS and FS PLMs, especially prompt-based generative large language models, offer viable and scalable alternatives to traditional supervised approaches. This study lays the foundation for the broader adoption of prompt-based PLMs in requirement engineering.

Read the paper · More papers on PaperTik