Replication Package for "Regression Test Prioritization Leveraging Source Code Similarity with Tree Kernels"
Francesco Altiero, Anna Corazza, Sergio Di Martino, Adriano Peron, Luigi Libero Lucio Starace · Zenodo (CERN European Organization for Nuclear Research) · 2023
Replication Package for the paper "Regression Test Prioritization Leveraging Source Code Similarity with Tree Kernels". The package contains both the dataset and the source code to replicate the experiments in the paper. The source code for the Java implementation of the experiment in the paper is in the Sources folder. All details to configure and compile the code can be found in the README.md file therein. The dataset is located in the Dataset folder. There are 3 subfolders within: Projects: all the original data for the benchmark projects in the paper. Inside each project's folder, there is a sorted_version.txt file which contains the SHA of the collected versions, ordered by the most recent version (top row) to the least recent (the bottom line), and the actual folder for each version. We included both the line granularity coverage reports (to build the fault matrices) and the method granularity level reports, which were actually used to perform the techniques. For sake of completeness, we included all the available versions for each project, even if they were not used in the experimentation due to the limited number of changes. The ordinal number related to each version in the paper refers to the (reversed) order the version appear in the sorted_version.txt file, counting from 1. FaulMatrices: the folder contains the fault matrices for each experimental pair, in CSV format. There is one subfolder for each project, and a nested subfolder for each version. In the version folder, there is one folder for each previous version used in the experimentation, which contains a CSV file for each variant, named by the number of the variant. The structure is schematized as / / / .csv. Each CSV file contains the name of test cases on the row and the id of the faults on the column. A cell has a value of 1 if the test on the row discovers the fault on the column, or 0 otherwise. Results: some pre-calculated results obtained by executing the experiments. The sqlite.db file contains all the results in a relational database, while the folder csv stores all the data within the database exported in CSV format and it was the base on which we obtained the Figures present in the paper.