Bug Evolution and Maintenance Effort in Deep Learning Frameworks: An Empirical Comparison with Non-Deep Learning Software Frameworks

Timilehin Ogundare, Maxime Lamothe · Zenodo (CERN European Organization for Nuclear Research) · 2026

RESEARCH SUMMARY This research investigate the bug frequency pattern in Deep Learning (DL) frameworks and use the non-DL frameworks for comparative analysis. Specifically, we mined over 500k issue reports from six popular DL frameworks, six popular non-DL frameworks, and 100 auxiliary frameworks across DL and non-DL python ecosystem from Githuib, which is their issue repository platform using the Github API. Next, we used a keyword search technique (from empirical evidence) to categorize the extracted issues into three categories namely: functional, non-functional, and non-functional-performance bugs using their title and description fields as search reference. We used the keyword search-approach because this same approach has been used in several prior studies (referenced in manuscript). This study hypothesize that there exists difference between DL and non-DL frameworks in bug frequency analysis. The answers to this hypothesis (accept or reject) is underscored by three research questions and backed by our research findings using our implementation codes and dataset uploaded in this repository. Notably, this research is motivated to understand how bug evolution differs per development stage in DL vs non-DL framework, and how these findings can benefit software engineering practitioners in the area of software quality assuranceto optimize current software and improve future software developments. This research is primarily quantitative as we only analyzed the bug reports that were submitted to GitHub. Follow the guide below to reproduce our findings: (A) Requirements: Install all libraries in requirements.txt Run the python script on any Python-enabled IDE of your choice such as IntelliJ All folders and structures should be maintained as it is in the uploaded files and extracted zip files (B) Dataset availability The zip file labeled "all_issues.zip" contains all the bugs in JSON format that were extracted across the main subjects frameworks and the auxiliary frameworks from GitHub platform. The zip file labeled "project_commits.zip" contains all the commits in JSON format that were extracted across the main subjects frameworks and the auxiliary frameworks from GitHub platform (C) Implementation Steps First, generate a GitHub access token to have call access to the GitHub API when using our data extraction Python scripts. Step 1: To extract issues from any repository on GitHub platform run the script extract_issue_from_github.py (insert GitHub access token and supply the script with list of repository and framework combination inside the file list_of_frameworks.txt in the format repository/framework on separate lines) Added note: you can extract pre-downloaded issues in the "all_issues.zip" file, which contains issues of the subject systems in JSON format Step 2: To extract commits from GitHub platform run the script extract_commit_from_github.py (insert GitHub access token and supply the script with list of repository and framework combination inside the file list_of_frameworks.txt in the format repository/framework on separate lines) Added note: you can extract pre-downloaded commits in the "project_commits.zip" file, which contains commits of the subject systems in JSON format Step 3: run the script "pre_process_data_into_dataframe_csv.py" to pre-process the JSON files of the extracted issues and commits by running the script Step 4: run the script "group_data_into_stages.py" to group dataset into categories andf timeframe (functional, non-functional and non-functional performance) Step 5: run the following analysis script: run "compute_median_stats.py" to compute the median statistics of different variables (such as duration between between bug submission and date closed). You can adjust it to be daily, weekly or monthly run "compute_kendall_statistics.py" to compute the Kendall statistics of analytical variables (such as distribution of daily buig submission) run "compute__data_distribution_analysis.py" to analyze the data distribution weights in different categories run "identify_famework_first_version_date.py" to identify the date the firtst version of the framework was submittedto GitHub using the GitHub API run "q1_framework_positioning.py" to identify the relative position of the DL and non-DL frameworks relative to the auxiliary frameworks run "ScottKnottESD_test.py" to compute the ScottKnottESD test of the distribution of issue submission run "extract_labeling_review.py" to extract and compute the Kappa inter-rater agreement based on the manual labeling experiment (D) General Comment: The file "Bug_Taxonomy_Rubrics.docx" contains the rubrics used by the independent labelers for the manual labeling task The file "Bug_labeling_result.csv" contains the classification result by the independent labelers for the manual labeling task Any folder that needs to be generated is handled by individual scripts and is auto-generated All graphical outputs can be found in the "outputs" folder We perform the ScottKnottESD test (ScottKnottESD_test.py) on Google colab platform as it contains processor chip needed to run the tool as defined by the tool owner. Otherwise, we ran all other scripts from our remote computer via intelliJ IDE environment.

Read the paper · More papers on PaperTik