Analysis of advanced aggregation techniques for software metrics
Bogdan Vasilescu · TU/e Research Portal · 2011
A popular approach to assessing software maintainability and predicting its evolution involves collecting and analyzing code metrics. However, metrics are usually defined on a micro-level (e.g., method, class), and should therefore be aggregated in order to provide insights in the evolution at the macro-level (system). In addition to traditional aggregation techniques such as the mean, median, or sum, econometric aggregation techniques on the one hand , such as the Gini, Theil, MLD, Kolm, Atkinson, and Hoover inequality indices, and threshold-based aggregation techniques on the other hand, such as the approaches in the quality models set forth by the SIG company or Squale consortium, have been proposed and applied to software metrics. In this thesis we study several main econometric and threshold-based aggregation techniques for code metrics, with the goal of distilling requirements for future aggregation techniques for software metrics. We perform our study along two directions. In the first part we assume a theoretical standpoint and study properties of these techniques relevant to aggregation of code metrics. Additionally, we show that root-cause analyses can be performed efficiently using the Theil index, and that the aggregation technique in Squale shares common properties with inequality indices, and has a formal relation to the Kolm index. In the second part we present the results of a series of empirical studies of all three categories of aggregation techniques considered (i.e., traditional, econometric, and threshold-based) using different metrics and aggregation levels. For example, we observe consistently high and statistically significant correlation between the Gini, Theil, MLD, Atkinson, and Hoover inequality indices for all metrics considered, i.e., aggregation values obtained using these techniques convey the same information regardless of the metric. However, even though it might seem that all these indices are equally appropriate, this is not true since different indices have different application domains, emphasize different dimensions of inequality and possess different decomposability properties. Finally, based on our theoretical and empirical analyses, in the third part we propose requirements for future aggregation techniques for software metrics.