ImprovMLCQ: A Feature-Enriched Dataset for Advancing Code Smell Detection

Joanne Carneiro, Jessica Ribas, Amanda Santana, Eduardo Figueiredo, Juliana Alves Pereira · 2025

Code smells are indicators of poor design choices in source code that negatively impact software quality. While manual detection of code smells is time-consuming, their automated detection requires high-quality datasets. This work evaluates an improved version of the dataset Madeyski Lewowski Code Quest (MLCQ), called ImprovMLCQ, which incorporates an extensive list of features extracted with four tools: CK, PMD, Organic, and Designite; along with several project characteristics. Our goal is to leverage these features to gain deeper insights into the detection or four code smells (Long Method, Feature Envy, Data Class, and Blob), assessing the effectiveness of different Machine Learning (ML) and Deep Learning (DL) models, and exploring the impact of feature selection on predictive performance. We evaluate fifteen ML algorithms and four DL algorithms using ImprovMLCQ, leveraging various feature engineering and selection mechanisms to optimize predictive performance. Our results show that the enriched dataset significantly boosts the performance of ML and DL models.

Read the paper · More papers on PaperTik