Embracing NLP-Enhanced Word Embeddings for Contextually Enriched Technical Debt Estimation in Software Code Comments: An Empirical Study

Ishpreet Kaur, Lov Kumar, Vikram Singh, Pratyush Mishra, Proksh Proksh · 2024

The evolution of software (SW), which emphasizes flexibility, speed, and responsiveness, coupled with exponential growth and the constant introduction of new technologies, has led to the amplification of various challenges due to faster releases, complex applications, and a lack of expertise. Managing and predicting technical debt (TD) is critical in the today's technology driven environment as it promotes the long-term sustainability and security of projects. Comments and discussions by the developers throughout the project development lifecycle, related to code commits that may cause TD, are referred to as Self-admitted Technical Debt (SATD). These can provide an accurate understanding of the characteristics of TD as well as the classification of it's many forms in the SW under review. However, manually identifying and analyzing such data has a variety of drawbacks and can be inaccurate. In this sense, this study utilizes natural language processing (NLP) enabled word embedding schemes and machine learning models to provide a framework for predicting different types of TD, e.g., code debt and architecture debt. To enable this, a novel pipeline is employed that includes feature selection using a variety of techniques, the use of SMOTE for class balancing, and a range of machine learning techniques. Empirical analysis of the developed models using benchmark metrics emphasizes the importance of class balancing for enhanced performance and suggests that treating all features as equally significant yields optimal results. Furthermore, through these metric scores, it is evident that the Extra Trees (EXTR) Classifier achieves the highest level of accuracy.

Read the paper · More papers on PaperTik