Payload-Aware Intrusion Detection with CMAE and Large Language Models

Yong Cheol Kim, Chanjae Lee, Young Yoon · ACM Transactions on Privacy and Security · 2025

Intrusion Detection Systems (IDS) play a vital role in network security, yet signature-based methods are limited by high false positive rates (FPR) and inability to detect novel threats. Recent AI-based approaches offer improved adaptability, but most rely on flow-level or statistical features, constraining their ability to analyze sophisticated payload-based attacks. To address these challenges, we present a dual-path IDS framework: Xavier-CMAE, a lightweight model using Hex2Int tokenization and Xavier initialization, achieves 99.9718% accuracy and a 0.0182% FPR without pre-training; and LLM-CMAE, which leverages pre-trained LLM tokenizers for enhanced detection, achieves 99.9696% accuracy and a 0.0194% FPR at higher computational cost. Experimental results on the CIC-IDS2017 dataset reveal a distinct trade-off between efficiency and Contextually Adept and Scalable (CAS) power, indicating that a modular approach may enable both real-time scalability and in-depth threat analysis. This work advances AI-powered intrusion detection by (1) introducing a modular, payload-centric dual-path architecture that combines lightweight and CAS detection for adaptive, layered security; (2) demonstrating that Xavier-CMAE achieves real-time scalability and state-of-the-art accuracy without embedding pre-training; and (3) exploring the effectiveness and future potential of integrating pre-trained LLM tokenizers for nuanced, selective threat analysis and robust IDS design.

Read the paper · More papers on PaperTik