A file recognition and classification scheme
Stephanie Duhon Crouch, Dick B. Simmons, Junkuk Kim · 1999
Vast quantities of software files may be stored on a computer system. Current file management tools organize files on a number of file attributes like file extension, file date, file name, and file size. Users can perform limited searches for files using some or all of these file attributes. There are many problems with the way files are currently managed and classified. The need for a new and improved file classification scheme will be established and a new file classification scheme will be introduced. The proposed classification scheme will provide complex searches on a number of file attributes. Searches can be expanded or narrowed down by using disjunctive or conjunctive logical operators. A number of file attributes are made available through the file management system. In addition, files can be classified by attributes that are not readily available to most computer users. These attributes include a number of software metrics that can be applied to files. Metrics can be applied to binary files or ASCII text files. In this dissertation the focus will be on applying metrics to ASCII text files, in particular text files containing programming language source code. Four experiments were conducted and the results demonstrated that metrics can be used to determine relationships between files, for example, a file being different version or duplicate of another file on the system. The results of the experiment also show that source code created to solve the same programming problem has a predictable range of reused lines. In addition, a model is demonstrated to predict if two files are versions of the same file or are duplicate files. A file gathering tool has been developed that (1) allows complex queries to be performed on a number of file attributes acquired through the file management system and file attributes that are generated for ASCII text files for programming source code, (2) can be used to classify files by relationships, such as one file is a version of another file, one file is a duplicate of another file, or a files is not related to any other files.