Authorship Attribution: A Principal Component and Linear Discriminant Analysis of the Consistent Programmer Hypothesis.

Jane Huffman Hayes · 2008

The consistent programmer hypothesis postulates that a feature or set of features exist that can be used to recognize the author of a given program. It further postulates that different test strategies work better for some programmers (or programming styles) than for others. For example, all-edges adequate tests may detect faults for programs written by Programmer A better than for those written by Programmer B. This has numerous useful applications: to help detect plagiarism/copyright violation of source code, to help improve the practical application of software testing, to identify the author of a subset of a large project’s code that requires maintenance, and to help pursue specific rogue programmers of malicious code and source code viruses. Previously, a small study was performed and supported this hypothesis. We present a predictive study that applies principal component analysis and factor analysis to further evaluate the hypothesis as well as to classify programs by author. This analysis resulted in five components explaining 96 % of variance for one dataset, four components explaining 92 % variance for a second dataset, and three components explaining 80 % variance for a third dataset. One of the components was very similar for all three datasets (understandability), two components were shared by the second and third datasets, and one component was shared by the first and second dataset. We were able to achieve 100 % accuracy of classification for one dataset, 93 % accuracy for the second dataset, and 61 % accuracy for the third dataset. Closer examination of the third dataset indicated that many of the programmers were very inexperienced. Consequently, two subsets of the programs were examined (the first written by programmers possessing a high level of experience and the second adding in less experienced programmers) and classification accuracy of 100 % and 89%, respectively, was

Read the paper · More papers on PaperTik