Applying Large Language Models to Issue Classification
Gabriel Aracena, Kyle Luster, Fabio Santos, Igor Steinmacher, Marco Aurélio Gerosa · 2024
Effective prioritization of issue reports in software engineering helps to optimize resource allocation and information recovery. However, manual issue classification is laborious and lacks scalability. As an alternative, many open source software (OSS) projects employ automated processes for this task, yet this relies on substantial datasets for adequate training. This research investigates an automated approach to issue classification based on Generative Pre-trained Transformers (GPT). By leveraging the capabilities of such models, we aim to develop a robust system for prioritizing issue reports accurately, mitigating the necessity for extensive training data while maintaining reliability. In our research, we have developed a GPT-based approach to label issues accurately with a reduced training dataset. By reducing reliance on massive data requirements and focusing on few-shot fine-tuning, we found a more accessible and efficient solution for issue classification. Our model predicted issue labels in individual projects up to 93.2% in precision, 95% in recall, and 89.3% in F1-score.