Automated fault localization : a statistical predicate analysis approach

Peifeng Hu · 2006

Fault localization is a major activity in software debugging. Previous related work shows that dynamic execution statistics of predicates are good indicators of faulty locations. Statistical fault localization approaches build statistical behavior models from aggregated execution data on program predicates and select the behaviors related with program failures. These approaches generally achieve more accurate results than other approaches. This thesis uses a general framework to describe and analyze statistical fault localization. It decomposes the process of statistical fault localization into discernible parts and studies independently various parts of the process. It identifies components crucial to the performance of the overall approach and points out the preferable directions from an analytical view. This thesis improves existing statistical fault localization approaches by proposing a simple and robust non-parametric hypothesis test to rank the fault relevance of the predicates at chosen program points. It also makes an important observation that faulty predicates tend to be clustered. Rather than simply mapping the ranking result to the faulty source locations, it proposes to locate faulty subpaths based on clusters of fault relevant predicates. Such a subpath helps maintainers understand how faults are formed and how their effects are propagated. The proposal is verified using the Siemens suite, which is also used by many representative software testing and debugging researchers in their experiments. Empirical results show that the ranking method outperforms the latest statistical fault localization techniques significantly in locating faulty statements. Furthermore, a significant portion of the faults are located on the faulty subpaths constructed by our approach. Thus, the approach is robust and promising.

Read the paper · More papers on PaperTik