Weaknesses in LLM-Generated Code for Embedded Systems Networking

Murray Dunne, Kylee Schram, Sebastian Fischmeister · 2024

Modern firmware development is done in a fast-paced, time-constrained environment. This pressure tempts developers to use generative AI to write code for them to save time. While this is a powerful tool with careful developer review, these reviews are commonly sacrificed to meet deadlines. This results in AI-written code existing verbatim, deployed in the firmware of devices finding their way into our cyber-physical environment. In the absence of developer oversight, we suggest that generative AI-written code does not sufficiently account for common software weaknesses. In this work, we explore a collection of modern Large Language Models (LLMs) and use them to generate code based on popular network standards. We fuzz this code to discover vulnerabilities in the code generated by the LLMs. We organize these vulnerabilities according to the Common Weakness Enumeration (CWE) and use them to develop a three-axis taxonomy of common LLM-generated weaknesses. Finally, we provide suggested input categories to more easily exploit these weaknesses in a black-box setting, as a first step towards fuzz testing for LLM-generated code in embedded systems networking.

Read the paper · More papers on PaperTik