Fuzzy Clustering and KNN Integration for Improved Software Effort Estimation in Mixed Data Scenarios
Jitendra Kumar Gardia, A. V. S. Pavan Kumar, Rakesh Nayak, Pushkar Dubey, Resham Lal Pradhan, Parul Dubey · 2025
The Software Development Effort Estimation (SDEE) is very crucial in the project management process as it is the basis of an accurate resource allocation and planning In SDEE, the high missing data has a major issue, therefore, the forecast precision and estimation will be unreliable. This study tackles the problem of incomplete datasets, through the application of an optimized imputation method on six publicly available datasets from the SDEE filed (COCOMO81, Kitchenham, ISBSG-R8, NASA, USP05, and USP05FT). The above datasets usually contain a heterogeneous mixture of both categorical and numerical attributes, demonstrating a real-life complexity. It uses a state-of-the-art fuzzy clustering based KNNI (FC-KNNI) method, optimized with particle swarm optimization (PSO) to fine tune clustering parameters and perform better on mixed datasets. This new line of reasoning preserves the uncertainty of categorical attributes through fuzzy logic to enhance the accuracy of imputation. In this research, performance was gauged based on the standardized accuracy (SA) and prediction accuracy (Pred (0.25)), and statistical significance testing was performed on the findings to verify these results. The performance of FC-KNNI yields 10% to 16% higher SA and 12% to 16% higher Pred (0.25) than classical KNNI on each dataset and the improvement was statistically significant (p < 0.05). The outcomes manifest that FC-KNNI can be considered as a worthy methodology for the imputation of missing data for SDEE.