AD ata Mining System for Investigation on Cause - Effect Relations in OSS
Manuel Piubelli, Barbara Russo · 2011
The availability of data on open source projects increased dramatically throughout the last decade. This tendency attracts the attention of many researchers who focus their attention on different software attributes and on how these affect each other. Commonly, such research involves the analysis of individual or small groups of projects and thus provides results that are strongly biased by the choice of these projects. To exploit the full value of the available information, it is necessary to consider at least a representative proportion of the universe of open source software. Such a universal set would allow us to draw generalizable conclusions about software development and the development process. These can both serve as a benchmark to re-evaluate current beliefs as well as build new knowledge valuable to both industry and academia. A quantitative investigation of open source software requires the automated collection and examination of large amounts of data. Such automation is fraught with problems, though. Project data such as source code, versioning history, bug data, and descriptive information is scattered around heterogeneous sources and thus differs in form and content. Moreover, much information originating from abandoned, incomplete, or small projects is inconsistent and can thus lead to unreliable results. Finally, it is difficult to support investigations concerning general software attributes as Size or Quality since such analyses require a high level of abstraction. In this thesis, we present OSSQuery, a data mining system that indexes, collects, cleans, updates, and analyzes publicly available project data to allow quantitative research on open source software. An efficient, extensible, and reliable design allows our system to construct a representative data set and to analyze a variety of questions on open source software on demand. In spite of primarily addressing cause – effect relations, the modular architecture can readily accommodate several types of analysis. OSSQuery indexes over 143,000 open source projects to date, it stores both historical and recent information on more than 2,500 projects. We performed a number of investigations to test our system. The outcomes confirm our primary hypothesis, namely that research on a limited set of projects might not be valid for all open source software. While our results commonly agree with outcomes of other studies, they also show that a significant proportion of projects do not fit these findings, yet.