AuthAttLyzer: A Robust defensive distillation-based Authorship Attribution framework
Abhishek Chopra, Nikhill Vombatkere, Arash Habibi Lashkari · 2022
Source Code Authorship Attribution (SCAA) is the technique to find the real author of source code in a corpus. Though it is a privacy threat to open-source programmers, it has shown to be significantly helpful in developing forensic-based applications such as ghostwriting detection, copyright dispute settlements, catching authors of malicious applications using source code, and other code analysis applications. Recent advances in SCAA techniques have performed exceptionally well on varied datasets. However, recent works on gradient-based attacks and universal perturbations can adversarially modify source code to reduce the accuracy of state-of-the-art classification techniques based on deep neural networks to as low as 10%. In this paper, we derive inspiration from recent advances in cyber-security and propose using the concept of defensive distillation to create a new architecture for source code authorship attribution with increased robustness against such adversarial attacks (reduce sent size). We empirically show that our approach, defensive distillation, and varied feature selection reduce miss-classification on perturbed source code files for Google code Jam and GitHub database while maintaining a 95% accuracy on legitimate source code files.