Network-based classification analysis of objects created by software engineering processes

Rodney Kent Madsen · 1993

This dissertation explores the problem of how to use classification analysis to enable empirically-guided software development. The first part describes a set of classification services intended to be useful across a wide variety of software engineering processes. The subsequent parts describe specific means for providing these classification services and the results of using them on data from real software engineering processes. Chapter 1 introduces classification analysis and describes the classification services. Chapter 2 introduces network-based classification models as a means of providing these classification services. The network-based models are applied to data from sixteen NASA projects in order to predict which components in a software system are likely to be high fault or high effort. The network models averaged 89.6% classification correctness, 69.1% completeness, and 79.5% consistency, an improvement over the results achieved by classification trees on the same data set. Chapter 3 demonstrates that network-based classification models can be optimized for classification completeness or consistency without unduly sacrificing correctness. When applied to the NASA data set, network models achieved a wide range of tradeoffs: from 83.9% completeness and 65.4% consistency to 31.5% completeness and 96.3% consistency. Chapter 4 describes how the completeness/consistency tradeoff capability of networks makes it possible to create hybrid empirical/analytical techniques for examining concurrent software. Classification analysis can focus static concurrency analysis techniques on the parts of a system where they are most likely to yield results. When applied to task interaction concurrency graphs for two different systems, network models were able to reduce the number of states in the graph while retaining a higher portion of states that lead to deadlock than were present initially. For three versions of the Dining Philosophers the network models were able to reduce the number of states in the corresponding task interaction concurrency graph by 85% while increasing by 300% the portion of states that lead to deadlock. For two analyses conducted on Chiron--an interface development and management system--the number of states was reduced by 75% while the portion of states leading to deadlock increased 10-30%.

Read the paper · More papers on PaperTik