Enhanced robustness in compilers with communicating stream x machine and machine learning for optimal anomaly detection, correction and testing

B. A. Sanusi · UWE Research Repository (UWE Bristol) · 2026

Compiler design plays an important role in ensuring that the translation of programs written in high-level language into executable code are correct. However, in todays’ safety-critical environments, security gaps, visible and hidden defects in compiler models are liability factors that needed to be addressed in the process of examining that a system meets specifications and requirements of its intended objectives. Nevertheless, the traditional compilers often struggle with detecting and correcting complex errors and adapting to evolving programming paradigms. These limitations originate from their dependence on static rule-based methods, which lack flexibility and intelligence. Hence, this research aimed at developing a novel approach to enhance compiler performance by integrating Machine Learning (ML) with the computational power of Communicating Stream X-Machine (CSXM) framework. The goal is to develop a more adaptable and intelligent compiler that effectively detect and correct compiler errors. In addition, this research developed a CSXM-based compiler and combines the CSXM’s formal methods with advanced ML techniques. Specifically, the machine learning models including Random Forest (RF), and Convolutional Neural Network (CNN) are trained on real-world compiler error datasets and integrated into the CSXM-based compiler to enhance its error detection, correction and optimization capabilities in the compilation process. Long Shot-TermMemory (LSTM) was used to optimize CNN by capturing temporal dependencies in the data. The performance of the ML-enhanced compiler was evaluated through correct testing and verification using accuracy, precision, recall, and f1-score across DeepFix dataset and Generated dataset. The CSXM theoretical technique was implemented in Visual Studio (2022) to develop the CSXM-based compiler.The DeepFix dataset results revealed that accuracy, precision, recall, and f1-score of RF algorithm were 85.51, 86.54, 84.11, and 85.31% respectively for the RF model classification phase. The corresponding values were 86.44, 85.46, 87.85, and 86.64% respectively for the CNN model classification phase. The improved model CNN-LSTM gave accuracy, precision, recall, and f1-score of 89.72, 89.00, 90.65, and 89.82% respectively. In addition, the Generated dataset results shows that accuracy, precision, recall, and f1-score of RF algorithm were 93.00, 91.30, 93.30, and 92.30% respectively for the RF model classification phase. The corresponding values were 91.00, 90.20, 92.00, and 91.10% respectively for the CNN model classification phase. The improved model CNN-LSTM gave an accuracy, precision, recall, and f1-score of 90.00, 88.90, 92.30, and 90.50% respectively. The CNN-LSTM maintains high accuracy across both datasets with minimal variance which provides stable performance and can further generalized well across similar datasets. This research contributes to the field of software engineering by introducing a novel hybrid approach to compiler design by integrating formal methods with data-driven ML techniques. The research advances the field by demonstrating the potential of ML to enhance compiler functionality by establishing a foundation for future exploration of AI-driven compilers and setting a new standard for intelligent, adaptable compilation process.

Read the paper · More papers on PaperTik