Test Script Repair of Deep Learning Library Testing
Xing Fu · 2025
Deep learning (DL) libraries such as TensorFlow and PyTorch are widely utilized to develop machine learning models. However, when testing DL libraries, evolving library versions and flaws of testing approaches may result in invalid models within test scripts, which influence testing efficiency severely. Repairing these models can be difficult to achieve manually because of the complexity of DL models. In this paper, we propose a novel approach that utilizes a specialized prompt-based strategy with Large Language Model (LLM) to repair invalid DL models. We manually provide structured error information and model configurations to LLM, allowing it to generate code to fix invalid models. Our work shows that most invalid models can be repaired successfully with our strategy. Moreover, our approach can help to detect flaws in the DL library testing approaches and issues caused by version updates, which enhances the robustness and transferability of DL models.