Identifying security fault reports via text mining
Michael Gegick, Pete Rotella, Tao Xie · NCSU Libraries Repository (North Carolina State University Libraries) · 2009
A fault-tracking (bug-tracking) system such as Bugzilla contains fault reports (FRs) collected from various sources such as development teams, test teams, and end-users.Software or security engineers manually analyze the FRs to label the subset of FRs that are security fault reports (SFRs), which indicate a security problem.These SFRs generally deserve higher priority in fault fixing than the not-security fault reports (NSFRs).However, this manual process is time consuming and error-prone (e.g.mislabeling an SFR as an NSFR).To address these important issues, we developed a new approach that applies text mining natural-language descriptions of FRs to train a statistical model on already manually-labeled FRs to identify unlabeled SFRs or SFRs that are manually-mislabeled as NSFRs.A security team can use the model to automate the classification of FRs for large fault databases to reduce the time that they spend on searching for SFRs.We evaluated the model's predictions on a large Cisco software system with over ten million source lines of code.Among a sample of FRs that Cisco software engineers manually labeled as NSFRs, our model successfully classified a high percentage (78%) of the SFRs as verified by a Cisco security team, and predicted their classification as SFRs with a probability of at least 0.98.Our results also indicate that a high percentage (77%) of the SFRs identified by our model is associated with software components that a code-level statistical model predicted to be attack-prone.Such findings provided valuable insights for calling for a future combined approach that exploits both textual information of FRs and code-level information of their associated software components.