Analyzing code fragments for faults
Ashwin Kallingal Joshy · 2023
It appears to be a fact of life that no matter how much effort is spent on testing a program, software defects are continually introduced and removed during the development process. Numerous static and dynamic code analysis based techniques exist to help the developers quickly locate and fix these faults. However, in spite of existence of advanced techniques like path-sensitive static analysis tools, fuzzers, automated fault localization, and program repairs, a significant amount of developer's time is still spent in locating, understanding and fixing faults. In this thesis, we developed code fragment analysis techniques to address these challenges. The code fragments consists of noncontinuous statements extracted from the original programs, with respect to some desired property. These code fragments can then be made executable to help gain further insight on the property. First, we identified a set of challenges for extracting and building such statements and devised solutions for them. Then, we used our findings to automatically validate path-sensitive static analysis warnings by converting the paths reported by the tools into code fragments. These code fragments are then dynamically executed to check for the existence of the static warnings reported by the tools. We found that code fragments are able to validate static warnings by exposing their dynamic symptoms. Hence, we created as-small-as-possible executable code fragments that can reproduce the faults called fault signatures. We found that the smaller size and complexity of the fault signatures helped fuzzers to generate crashing inputs for faults that are harder to reproduce using the original programs. As the next step, we generated fault signatures for crashes reported by fuzzers. We then used them to identify "unique" faults and deduplicate the crashing inputs reported by the fuzzers. We found that using fault signatures correctly grouped the crashing inputs for unique faults and outperformed the state-of-the-art fuzzer based deduplication methods by as much as 75 times. Finally, we analyzed features of code fragments and used diffs generated from vulnerability fixing patches to actively train a machine learning model that identifies statements that contribute towards the bug fix. This model was then used to produce a large dataset of line-level vulnerability labels that can be used by other deep-learning based approaches for tasks like identifying vulnerable lines, bug detection, and fault localization. We found that using our dataset improved the performance of LineVul, a state-of-the-art vulnerability detection model, in terms of both the number of detected vulnerable functions/statements and their accuracy.