A Contrastive Learning Approach to Bug Severity Classification with Large Language Model Embeddings
Mosarrat Rumman, Emon Roy, Anushka Zaman, Jeremy S. Bradbury · 2025
Automatically classifying bug severity helps reduce manual effort and improve response times in software maintenance. This study leverages Large Language Models (LLMs), specifically CodeBERT, to generate contextual embeddings of bug reports for automated severity classification. To enhance the quality of embeddings, we integrate Contrastive Learning, which structures the embedding space by bringing similar bug reports closer and pushing dissimilar reports apart. Our approach is evaluated on the NASA PITS and Mozilla datasets and compared against LLMs fine-tuned without contrastive learning as well as traditional embedding models like Doc2Vec. Our results demonstrate that contrastive learning consistently enhances performance, particularly on imbalanced and diverse datasets. Furthermore, the results demonstrate that LLMs excel in handling longer bug descriptions, while traditional embedding models like Doc2Vec are suitable for smaller, structured datasets.