Balancing Validity and Vulnerability: Knowledge-Driven Seed Generation via LLMs for Deep Learning Library Fuzzing
Rongtao Liao, Xuehu Yan, Zeshan Pang, Kai-Long Zhu · Applied Sciences · 2025
Fuzzing deep learning (DL) libraries is essential for uncovering security vulnerabilities in AI systems. Existing approaches enhance large language models (LLMs) with external knowledge such as bug reports to improve the quality of generated seeds. However, most approaches still rely on static strategies or single knowledge sources, limiting their ability to produce syntactically valid inputs that also expose deeper bugs. To address this challenge, we propose an adaptive seed generation approach that models knowledge-guided prompt selection as a multi-armed bandit problem. Our method first constructs two knowledge bases from API documentation and bug reports, then dynamically selects and refines prompt strategies based on real-time feedback. These strategies are tailored to the knowledge types in the respective bases. We design a multi-dimensional reward function to evaluate each batch of generated seeds by measuring their error-triggering potential and behavioral diversity, enabling a balanced exploration of both syntactically valid and bug-triggering test cases. Our experiments on three DL libraries, PaddlePaddle, MindSpore, and OneFlow, identify 17 previously unknown crash bugs, demonstrating the effectiveness and generalizability of the proposed approach.