Learning the Structure of Commands by Retraining a Language Model
Zafar Hussain, Lalli Myllyaho, Jukka K. Nurminen · 2024
In the field of cybersecurity, learning the command-line commands’ syntax holds paramount importance in distinguishing valid and malicious commands. To learn the syntax of command-line commands, we curated an extensive dataset of Windows 10 command-line commands, developed a specialized vocabulary, and trained a custom tokenizer equipped with a masked language model head. Comparative analyses against traditional methods, including a second-order Markov Model and a Regular Expression-based system, unequivocally demonstrated the language model’s superior proficiency. Employing clustering algorithms like DBSCAN, HDBSCAN, and OPTICS allowed us to categorize command-line commands based on their syntactical similarities, revealing the model’s excellence in understanding sequences and detecting syntax with minimal noise. Manual analyses of command syntax, complemented by BERTScore assessments, consistently yielded metrics exceeding 0.90 for precision, recall, and F1-score. These robust results affirm the model’s high accuracy and effectiveness in learning command syntax. In conclusion, our language model not only helps in enhancing protective measures against malicious activities but also showcases adaptability to the ever-evolving nature of command-line commands’ syntax.