PEARL: An Adaptive and Explainable Hardware Trojan Detection Using Open Source and Enterprise Large Language Models
Ripan Kumar Kundu, Khurram Khalil, Eric P. Garcia, Ethan Grassia, Prasad P. Calyam, Khaza Anuarul Hoque · IEEE Access · 2025
The Integrated Circuit (IC) supply chain risk allows attackers to implant hardware Trojans (HT) in various stages of chip production. To counter this, different machine learning (ML) and deep learning (DL)-based methods have been developed to detect HTs. However, these methods require massive amounts of high-quality labeled data for effective training, extended training times for accurate HT detection, limited generalization to novel or unseen HTs, and insufficient capability to explain the detected HTs. Recent studies have started exploring the potential of Large language models (LLMs) for hardware security tasks. However, there are no current studies that explore, study, and compare the applicability of open-source vs. enterprise LLMs for efficient HT detection and explanation of the detected HT. To close the gap, we propose an innovative HT detection and explanation method by leveraging the knowledge of a pre-trained LLM, namely enterprise application programming interface (API)-enabled a Generative Pre-trained Transformer (GPT)-3.5 Turbo, Google Gemini 1.5 Pro and open-source LLMs: Meta AI Llama-3.1 and DeepSeek AI DeepSeek-V2 models, which have already been trained on massive and diverse datasets and is capable of providing the reasoning of the detected HT. Specifically, we apply In-Context Learning (ICL)-based mechanisms: zero-shot, one-shot, and few-shot learning strategies (e.g., register transfer level (RTL) files (Verilog) of the circuit) to adopt this model for HT detection and explanation tasks. We validate our proposed approach on diverse circuit design benchmarks from Trust-Hub and ISCAS (85 and 89). Our experimental results show that the proposed few-shot learning-based enterprise-API-enabled GPT-3.5 Turbo and open-source DeepSeek-V2 LLM models detect unknown HTs with an accuracy of (97% and 91% ) and drastically reduce the training time compared to state-of-the-art techniques. Furthermore, after detecting the HT, they provide human-centric reasoning/explanation, reinforcing transparency and trust in the IC supply chain through its understanding.