Adversarial Label Flipping Attack on Supervised Machine Learning-Based HT Detection Systems
Richa Sharma, G. K. Sharma, Manisha Pattanaik · 2024
In the semiconductor landscape, safeguarding integrated circuits against Hardware Trojans (HT) is critical.To address this pressing concern, supervised machine learning (ML) has emerged as a promising defense mechanism for HT detection. However, the vulnerability of supervised ML-based defense mechanisms to adversarial attacks poses a substantial challenge, potentially compromising model prediction performance. This paper presents a label flipping poisoning attack, strategically targeting supervised ML-based HT detection systems during the pre-silicon IC design phase. Leveraging the power of the Isolation Forest, our method first identifies potential Trojan nets through a process of random partitioning, flips their labels, and perturbs the model training process. Further, random subsampling is applied to select a subset of Trojan-free samples, whose labels are also flipped. This model-independent, untargeted, and black-box attack is evaluated against Trust-Hub & DeTrust Benchmarks, resulting in a substantial reduction in model recall, challenging the reliability of ML-based HT detection systems.