Software Artifact Mining in Software Engineering Conferences: A Meta-Analysis - Replication Package
Zeinab Abou Khalil, Stefano Zacchiroli · Zenodo (CERN European Organization for Nuclear Research) · 2022
Software Artifact Mining in Software Engineering Conferences: A Meta-Analysis - Replication Package Contents This is a replication package for the paper entitled ["A Systematic Mapping of Software Artifact Mining in Software Engineering Conferences"]. It contains 16 years of history for each of the following conferences: ICSE, International Conference on Software Engineering ASE, IEEE/ACM International Conference on Automated Software Engineering FSE, ACM SIGSOFT Symposium on the Foundations of Software Engineering ICSM, IEEE International Conference on Software Maintenance MSR, Working Conference on Mining Software Repositories WCRE, Working Conference on Reverse Engineering ICSME, International Conference on Software Maintenance and Evolution ICPC, IEEE International Conference on Program Comprehension SCAM, International Working Conference on Source Code Analysis & Manipulation The data is stored in a PostgreSQL database (see the [PostgreSQL Dump] in folder db). Alternatively, the database can be recreated from CSV files using Python and the SQLAlchemy Object Relational Mapper using the scripts included (more details below). Data Papers and authors: the DBLP data dump. We used the data in dblp-2021-11-02.xml file. Using the replication Directly Most simply, you can import the [SQL dump] in the folder db into your database management system and start querying. Via Python Alternatively, you can take a look at how the database was created using PostgreSQL, Python and SQLAlchemy, and use these mechanisms also for querying. This will allow you to easily extend the database or update its schema. Dependencies and installation instructions If you take this path, make sure you have Python and a PostgreSQL server installed before attempting anything. Follow the follwoing steps (tested on our OS 11.3 machine with Python 3.7.7): Install SQLAlchemy: easy_install SQLAlchemy Tweek database.ini for your particular PostgreSQL user and password (the script assumes user root with empty password) Install CERMINE [https://github.com/CeON/CERMINE] to extract content from PDF files or use cermine.jar file included here. Python scripts initDB.py: declares the database schema using Python classes (will be automatically mapped to tables by SQLAlchemy). populateDB.py: reads data about the papers for each conference and loads it into the database. downloadPdf.py: download the pdf of the papers using a modified version PyPaperBot (The source code of our PyPaperBot is in the replication package). cermine.py: Extract the text from the Pdf files into XML files the pdf. How to use Python files arguments: Arguments Description Type --dir Directory path in which to save the result (str) Jupyter notebooks The various Jupyter notebooks containing all the scripts used to gather the data and answer the research questions. 1_main.ipynb: Runs populateDB, downloadPdf and cermine files. 2_SectionHeadersExtraction.ipynb: Extract the header of the section which helps us exclude irrelevant sections. 3_NLP.ipynb: Generate n-grams and update the database. 4_PreliminaryAnalysis.ipynb: 5_RQ1-Artifacts.ipynb: Execute the code and generate the figures to answer RQ1. 6_RQ2-ArtifactCombinations.ipynb: Execute the code and generate the figures to answer RQ2. 7_RQ3_Purposes.ipynb: Execute the code and generate the figures to answer RQ3. db : a repository containing all the data of our study after all processing steps. figs : The generated figures from the notebooks for the paper.