Causal Tracing to Identify Hacking Knowledge in Large Language Models
Ryan Marinelli · 2024
This research seeks to identify where hacking knowledge is stored in models. SQL Injection is used as a base case considering the risks involved with production systems. Through using causal tracing techniques, it is determined that SQL Injection knowledge capabilities are most prevalent in the last layers of the network. It is also established that there may be some shared representation of SQL Injection knowledge across models through training a KNN classifier. When training the classifier on GPT-Neo’s activations and attempting to classify on GPT-2, it was found to be 52% accurate.