Enabling Early Lifecycle Predictive Models of Software Systems
Rick Selby, Jairus Hihn · 2006
Predicting the outcomes of software projects using data available in early development phases allows developers and managers more time to re-plan or re-design systems as needed to increase success. This research investigates early lifecycle predictive models of software faults and effort using information available no later than the system design phase. Static analysis of system designs characterizes intra- and inter-component control flow structure to calculate objective system design metrics. The model building process composes these design metrics into a set of relative factors that calibrate the relationships between the metrics and software faults and effort. These component-based predictive models provide a powerful approach to capture and analyze the designs of complex systems. These models facilitate the identification of essential system characteristics in terms of component structures and interactions, and they enable comprehensive analysis by being scalable to large systems. The research approach applies static analysis tools to the designs of 23 large-scale NASA systems containing over 5000 software components in order to automatically generate state- and interaction-based models of their designs. The static analysis tools derive the number of internal states for a component from its internal control flow design structure using cyclomatic complexity, and they calculate the number of interactions between components from the passing of control flow or data abstractions between components. The research approach evaluates the models by characterizing their relationships to favorable or unfavorable aspects of the system designs, in terms of component development faults and effort. The data analysis focuses on describing the characteristics of the state- and interaction-based models and identifying system design factors revealed by the models that relate to low or high faults, fault correction effort, and development effort. The results include the following, with comparative data being statistically significant at α < 0.05: ♦ The static analysis tools successfully generated state- and interaction-based models of the 23 system designs comprised of 5469 components. The numbers of states per component were exponentially distributed with an average of 14.3 and a median of 7.0. The numbers of interactions per component were exponentially distributed with an average of 8.8 and a median of 4.0. During system development, the components had 0.60 faults on average. ♦ Components with the most states (top 20%) had 7.8 times more faults than did those with the fewest states (bottom 20%). Components with the most states also had 22.5 times more fault correction effort and 5.4 times more development effort than did those with the fewest states. ♦ Components with the most interactions (top 20%) had 2.9 times more faults than did those with the fewest interactions (bottom 20%). Components with the most interactions also had 3.9 times more fault correction effort and 2.7 times more development effort than did those with the fewest interactions. In conclusion, we outline future research directions that build on these system modeling ideas and strategies.