HADA: Leveraging Multi-Source Data to Train Large Language Models for Hardware Security Assertion Generation

Weimin Fu, Yiting Wang, Zelin Lu, Xiaolong Guo, Gang Qu · 2025

Hardware security verification is critical but labor-intensive, particularly in crafting security-focused assertions for assertion-based verification. While large language models offer automation potential, their performance is constrained by limited domain-specific training data and fragmented hardware security knowledge. We present HADA (Hardware Assertion through Data Augmentation), a novel framework that systematically constructs a fine-tuning dataset for LLMs using three complementary sources: (1) formal verification outputs, (2) hardware vulnerability databases, and (3) version control histories of real-world designs. This multi-source data augmentation enables the generation of high-quality SystemVerilog assertions and their contextual explanations. We fine-tune several open-source LLMs and evaluate them on HSAEval, a new benchmark for hardware security assertion tasks. Results show that HADA significantly improves assertion correctness and coverage, outperforming prior syntactic and functional verification approaches. All models, datasets, and raw evaluations will be released to support open research in LLM-driven hardware security verification.

Read the paper · More papers on PaperTik