Detection of Fault-Prone Classes Using Logistic Regression Based Object-Oriented Metrics Thresholds

Shahid Hussain, Jacky Keung, Arif Ali Khan, Kwabena Ebo Bennin · 2016

Background: In the plethora of studies, the object-orientedmetrics have been empirically validated to assess the design properties and quantify the high-level quality attributes such as fault-proneness, either at the method or class granularity levels of software. Motivation: A more precise value of an object-oriented metric can be used as an indicator for the developers tomake the informed decisions regarding the detection of design flaws and classify the fault-proneness classes. Method: Bender used an approach in the domain of epidemiology studies to derivethe threshold values for the risk factors. In our study, we follow the Bender's approach and propose a model to derive the thresholds for a set of software design metrics via non-linearfunctions, which are described through logistic regressioncoefficients. Subsequently, we perform four types of analysis and three experiments in order to evaluate and compare the effectiveness of derived thresholds in the domain of classificationof fault proneness classes. We use the Precision, Recall, Fmeasureand classification accuracy performance measures toassess the effectiveness of derived metrics thresholds. Results: We compare the derive threshold values of DIT, CA, LCOM andNPM metrics with their existing data distribution basedthreshold values, and observed the significant increase in the classification accuracy of fault-prone classes. For example, DIT(27%), Ca (2%), NPM (2%) and LCOM (15%) for the Ant-1.5project. Conclusion: The analysis results suggest that the proposed model can be applied to derive the thresholds of otherobject-oriented metrics which present either with or withoutheavy-tailed distribution, however, the proposed model to derivethresholds cannot generalize for all the systems due to variationin data characteristics.

Read the paper · More papers on PaperTik