Foundation Models for Automatic Issue Labeling
Giuseppe Colavito · 2025
Foundation models are transforming software engineering practices through their ability to understand and generate code, process natural language, and automate various development tasks. Despite their potential, effectively applying these models to specialized software engineering tasks remains challenging due to the need for domain-specific understanding and accurate labeling of data. This research project investigates how foundation models can be leveraged to automate labeling tasks in software engineering, with a specific focus on issue classification as a representative case study. Issue tracking systems, while essential for collaborative software development, often suffer from misclassification problems that require significant manual effort to correct. We explore how foundation models can be adapted to automatically label issues accurately, reducing the need for manual intervention while maintaining high-quality classification. The project examines several key aspects: the capabilities of different foundation models in understanding software engineering artifacts, methods for adapting these models to specific labeling tasks through techniques like prompt engineering and few-shot learning, and approaches for integrating automated labeling into real-world scenarios. This research contributes to the broader understanding of how foundation models can be effectively applied to reduce manual labeling efforts across various software engineering contexts, using issue classification as a concrete demonstration of their potential.