Improving Software Reliability Through Bug Detection and Automated Error Repair Using Transformer-Based Models
Nedal Mohammed Nwasra, Jamal Zaraqu, Zahid Hussain Qaisar · 2025
Automated bug detection and repair are critical in determining the reliability and cost of software development. To address these concerns, this research uses a transformer-based sequence-to-sequence model called CodeT5, together with the Defects4J dataset, which contains real Java programs with officially identified buggy and fixed versions. The approach involves retrieving buggy and fixed code pairs, pretraining CodeT5 for sequence-to-sequence learning, and then evaluating the model using the BLEU, EM, and ED measures. This model got translated sentences' BLEU 78.4%, Exact Match, rate 64.2% and accuracy 82.0%. While having an average edit distance of 12.3 operations, the generated fixes suggest how slight modifications can be made to produce semantically and syntactically correct fixes with little human intuition or input. This work shows that program repair can benefit from applying state-of-the-art deep learning algorithms and realistic benchmark datasets. The outcomes point to possible uses for enhancing directions of software creation and shortening the hours spent on debugging. Future work will be dedicated to using similar techniques in different languages and handling different types of bugs.