CodeComClassify: Automating Code Comments Classification using BERT-Based Language Models

Khubaib Amjad Alam, Wajid Ali, Summan Aziz, Muhammad Haroon, Meer Hashaam Khan, Zahoor Ahmad, Nadeem Abbas · 2025

Code comments are pivotal in enhancing code read-ability, maintainability, and team work in software development. As volume of code comments escalates in large software projects, it becomes impractical to manage and understand code comments, which prompts the need of exploring automated approaches for reliable handling of code comments. This paper introduces CodeComClassify, an automated code comments classification approach which is built on the Transformer-based pretrained model, distilbert-base-uncased model. CodeComClassify proficiently classifies code comments into 19 distinct categories across three languages: Java (7 categories), Python (5 categories), and Pharo (7 categories). The workflow involves cleaning and preprocessing datasets provided by NLBSE'24 Tool Competition, hyperparameter tuning, fine tuning pretrained Transformer model, distilbert-base-uncased, and training a custom multi-label model for each language. Multi-label models for Java, Python, and Pharo are trained and evaluated on 10,555, 2,555 and 1,765 class-level code comments, respectively, extracted from 20 open source projects. The results indicate that distilbert-base-uncased model demonstrates a promising level of performance by achieving a 81% average F1-score (weighted average).

Read the paper · More papers on PaperTik