The Classification of Software Requirements using Large Language Models: Emerging Results and Future Directions
Yizhuo Zhang, Yanhui Li · 2024
In recent years, large language models (LLMs) have demonstrated exceptional capabilities across various tasks. However, the effectiveness of LLMs in software requirement classification has yet to be thoroughly explored. This paper aims to bridge this gap by evaluating the performance of multiple LLMs (including GPT-3.5, GPT-4, GPT-4o, HuggingChat, Gemini, and DeepSeek) under different task prompts, focusing specifically on the binary classification of functional requirements (FR) versus non-functional requirements (NFR), as well as the multi-class classification of non-functional requirements. Experimental results show that HuggingChat achieved the highest accuracy and F1 score in the binary classification task, outperforming other models; meanwhile, GPT-4 performed relatively well in the multi-class classification task. Through the analysis of these results, this paper provides important references for future research and applications in software requirement classification and explores potential directions for LLM-based software requirement classification.