An empirical evaluation of the effectiveness of various ML, DL, and CodeBERT models to enhance the quality of software with the application of AST and Embedding techniques
Sonika Chandrakant Rathi, Lalita Bhanu Murthy Neti, Lov Kumar · 2025
The competitive market, dynamic business requirements and rapid growth in technologies poses a lot of challenges to deliver quality software.The application of Data Mining (DM), Machine Learning (ML), and Deep Learning (DL) for the software engineering activities such as software fault prediction, software maintainability prediction, fault localization, code refactoring and cloning etc, improves the quality of software, expedites its development and enhances the productivity of developers.On top of the existing static features and classical learning models, the recent release of CodeBERT and state-of-art deep learning models offer a promising hope of handling a wide range of software engineering operations efficiently to improvise the software quality.The proposed research topic focuses on Software Fault Prediction (SFP), and Fault Localization, two software engineering tasks to further enhance the quality of software systems through the application of DM, ML, and DL techniques.The performance of the Software Fault Prediction (SFP) model is hampered by feature redundancy, correlation, and irrelevance.The application of a prediction model to such an imbalanced class or software source code metric yields incorrect prediction results.In addition, the performance of the SFP model differs depending on which ML methods and approaches were used to train it.Hence, an extensive study of learning models with possible changes in the associated techniques along with detailed empirical results is necessary to find the most effective and high-performing SFP model.To establish the cost-effectiveness of SFP models, a costbenefit analysis of the applied ensemble approaches is also required.By analysing a variety of dynamic execution information (such as bug reports, test results, and failed/passed tests), fault localization provides developers with the ability to locate potentially faulty code files and preferably, if possible, localize them to segments of code or methods.The fault localization task has been studied in the past using a variety of approaches, including those based on information retrieval (IR), spectral analysis, and learning-based.The challenge of fault localization techniques is test cases for spectra-based methods are rarely available and in most cases, IR-based code only functions