Toward Hardware Security Benchmarking of LLMs
Raheel Afsharmazayejani, Mohammad Moradi Shahmiri, Parker Link, Hammond A. Pearce, Benjamin Tan · 2024
With the rapid advancement and proliferation of large language models (LLMs), there is a pressing need to explore and, crucially, evaluate their utility. Recently, LLMs have shown promise in digital design, with evidence of some ability to produce functional HDL code. However, to better understand LLM capabilities and guide the ongoing development of LLMs, we need approaches to evaluate the quality of generated artifacts across myriad dimensions. Thus, this work proposes an approach for evaluating the security of LLM-generated designs, which is especially important as security is an ongoing concern. We provide new insights into the challenges and desiderata for benchmarking LLMs for hardware security risks. This paper outlines our initial work developing a security-focused evaluation suite for LLM-aided HDL generation. We present an illustrative preliminary use of our evaluation suite to show the insights we can gain from security evaluation.