An Empirical Study of Data Mining Code Defect Patterns in Large Software Repositories
Kingsum Chow, Zhongming Wu, Xuezhi Xing, Zhidong Yu · 2009
There has been a growing interest in mining software code defect patterns and using this knowledge to identify potential problems [ 2, 8- 15]. To understand the benefits of such methods, we applied them to several large software repositories. We learned the effectiveness and the limitations of applying these methods. These methods are called “static analysis” as they analyze the code for defects without execution. They are different from the traditional static analysis as they apply data mining in the analysis, while the traditional static analysis uses program analysis, such as data flow analysis and control flow analysis. There is a trend to combine the data mining methods and program analysis techniques to gain more effective results. Many tools are available to detect software defects. But if the tools have no knowledge about what to check they can’t find defects [ 5]. There are common defect patterns such as buffer overflow, null pointer dereference and several tools address these patterns such as Purify [ 19], FindBugs [ 20]. However, application specific defects can be difficult to find, probably because there are few common patterns among different applications. Therefore some approaches, such as DynaMine [ 2], try to explore the patterns of specific applications automatically, while others try to provide description based methods to describe these patterns as assertion statements. Examples of tools and methods that fall into this later category are contract programming, AOP (Aspect Oriented Programming), PQL [ 3], and Metal [ 4]. The automatic approaches such as DynaMine may have difficulty in offering a good solution to explore complex patterns. The description based approaches such as PQL are often powerful at describing patterns, but they require manually constructing the specifications. Such tasks can be overwhelming [ 5]. In this paper, we dive into industry-level projects such as Harmony [ 16] an open source virtual machine to analyze and summarize their defect patterns using different pattern analysis methods. We gain some insights into the classifications of common code defect patterns. We make use of data mining methods to extract usage patterns automatically from the source code of different projects. We also look into the issues tracking system and revision history to analyze and summarize patterns. We will apply these tools and methods to detect pattern-related defects in new source code as well. Several methods are evaluated for their effectiveness to detect defect patterns. The contributions of this paper are: (1) an empirical study to evaluate the effectiveness of data mining software code defects using a few large software repositories, and (2) insights into the characteristics of the software systems and the common code defect patterns.