Towards Next-Generation Methods for Hardware Security Assurance
Mohamadreza Rostami · TUbilio (Technical University of Darmstadt) · 2026
Computing systems have increasingly relied on sophisticated hardware architectures to meet the growing demands of modern applications. High-performance processors employ complex microarchitectural optimizations such as speculative execution, out-of-order execution, and deep memory hierarchies to maximize efficiency. While these mechanisms significantly improve performance, they also expand the attack surface and introduce critical security vulnerabilities. In recent years, attacks such as Spectre, Meltdown, Zenbleed, and numerous other microarchitectural exploits have demonstrated that vulnerabilities in hardware can undermine the security guarantees of the entire software stack. However, hardware vulnerabilities are particularly difficult to mitigate once deployed, as hardware designs cannot be easily patched or updated after fabrication. Consequently, ensuring the security of hardware designs has become a critical challenge for modern computing infrastructure. Traditional hardware security assurance techniques, including static analysis, simulation-based testing, and formal verification, play an essential role in validating functional correctness. However, these approaches often struggle to detect complex security vulnerabilities arising from interactions among microarchitectural components. As a result, hardware fuzzing has recently emerged as a promising methodology for automated hardware security evaluation. Inspired by the success of software fuzzing, hardware fuzzing algorithmically explores the behavior of hardware designs by generating large numbers of test inputs and monitoring the execution for unexpected behavior. Despite its advantages over traditional hardware security assurance techniques, existing hardware fuzzing approaches still face several fundamental challenges, including limited ability to automatically detect microarchitectural vulnerabilities, difficulties in generating semantically meaningful test inputs, inefficiencies due to expensive hardware simulations, and limited applicability to post-silicon systems deployed in real-world environments. In this dissertation, we design, implement, and evaluate novel techniques that significantly advance hardware fuzzing toward a scalable, intelligent security analysis methodology. In particular, we introduce new approaches for (1) autonomous detection of microarchitectural vulnerabilities, (2) generation of semantically meaningful test inputs using Artificial Intelligence (AI) techniques for general hardware fuzzing, (3) improving the efficiency and scalability of the hardware fuzzing method, and (4) extending hardware fuzzing to proprietary post-silicon black-box hardware, providing a completely new perspective to fuzzing unknown hardware designs. Autonomous Detection of Microarchitectural Vulnerabilities. We introduce two novel hardware fuzzing frameworks, Specure and WhisperFuzz, designed to detect and localize microarchitectural vulnerabilities in modern processors. In Specure, we propose a hybrid speculative leakage detection framework that combines hardware fuzzing with Information Flow Tracking (IFT) to automatically identify speculative execution leakages without relying on predefined attack templates or Golden Reference Model (GRM). Specure further introduces a leakage path coverage metric that models the speculative leakage search space and guides the fuzzer toward unexplored microarchitectural behaviors. Complementing this approach, WhisperFuzz introduces a hardware fuzzing framework for detecting timing side-channel vulnerabilities directly at the Register-Transfer Level (RTL) level. The WhisperFuzz framework extracts microarchitectural state transitions from hardware designs and introduces a timing-aware coverage metric that enables systematic exploration of timing behaviors and automated localization of vulnerability root causes. Generation of Semantically Meaningful Test Inputs using AI Techniques. To overcome the limitations of traditional input generation, we explore the use of AI techniques for hardware fuzzing. Much like natural language, instruction sequences rely on structured semantics and dependencies, and naïve generation that ignores these relationships produces largely ineffective test cases. One natural approach is to leverage AI to generate structured instruction sequences. However, directly applying off-the-shelf Large Language Model (LLM)s is insufficient, as they do not adequately capture the domain-specific structure and dependencies required for effective hardware exploration. To address this, ChatFuzz leverages specialized LLMs with custom tokenization and training on machine code to generate semantically meaningful instruction sequences that can trigger complex processor behaviors. By incorporating hardware coverage feedback into the training process, ChatFuzz enables targeted exploration of microarchitectural state spaces. Building on this concept, GenHuzz introduces a Hardware-Guided Reinforcement Learning (HGRL) framework in which the processor itself acts as a teacher that continuously refines the fuzzing policy. This approach enables the generation of structurally coherent instruction sequences and significantly improves coverage growth and vulnerability discovery compared to existing fuzzing techniques. Improving the Efficiency and Scalability of Hardware Fuzzing. Hardware fuzzing faces fundamental scalability challenges due to the high computational cost of RTL simulation. To address this limitation, we propose two techniques that significantly improve fuzzing efficiency by leveraging the concept of a digital twin of the hardware. The HFL introduces a coverage prediction model that estimates the impact of generated test inputs before executing them on the hardware model, thereby reducing unnecessary simulations. GoldenFuzz complements this approach by introducing a Golden Reference Model (GRM)-guided fuzzing architecture that performs large portions of test-case refinement using a fast Instruction Set Architecture (ISA)-level reference model as a digital proxy, before executing promising candidates on the hardware design. Together, these techniques reduce simulation overhead and enable deeper exploration of complex processor architectures. Extending Hardware Fuzzing to Proprietary Post-silicon Hardware. Finally, this dissertation extends hardware fuzzing to real-world processors deployed in modern computing systems. We introduce Fuzzilicon, the first gray-box fuzzing framework that enables microarchitectural feedback-guided fuzzing of proprietary x86 processors at the post-silicon stage. Fuzzilicon leverages the processor’s microcode patching mechanism to obtain fine-grained microarchitectural-level coverage feedback without requiring access to RTL designs. In combination with a lightweight bare-metal hypervisor infrastructure and differential vulnerability detection mechanisms, this approach enables exploration of internal processor behavior and scalable vulnerability finding on real silicon. Taken together, these contributions transform hardware fuzzing from a largely heuristic, manually guided process into an algorithmic, intelligent, and scalable methodology for hardware security assurance. By enabling autonomous microarchitectural vulnerability detection with precise localization, improving the semantic quality of generated test inputs, and significantly reducing the cost of hardware exploration, this work advances the state of the art in both effectiveness and efficiency of hardware fuzzing. Moreover, by extending these techniques to post-silicon processors and proprietary architectures, this dissertation bridges the gap between academic research and real-world deployment. The presented approaches not only identify known vulnerabilities with substantial speedups but also uncover previously unknown security flaws in processor designs. Overall, this work establishes a foundation for next-generation hardware security assurance frameworks that are capable of algorithmically analyzing increasingly complex computing systems. Beyond academia, these contributions have influenced industrial practices, including their integration into security assurance pipelines within major semiconductor companies such as Intel. In addition, the results of this thesis have significantly contributed to ongoing efforts in hardware security standardization through collaboration with the Common Weakness Enumeration (CWE) program at MITRE Corporation, a U.S.-based not-for-profit organization that develops and maintains widely used cybersecurity standards such as CWE and CVE. Furthermore, this work has supported the objectives of the European Research Council (ERC) Advanced Grant HYDRANOS, contributing to multiple research pillars.