Phase-Adaptive Model Routing in LLM-Driven SSH Honeypots: Balancing Response Fidelity and Latency Across the Attack Lifecycle
Raiymbek Magazov, Fatima Uralova, Kuanysh Abeshev, Гүлдана Ахмеди, Gulnur Aksholak · Future Internet · 2026
Secure Shell (SSH) intrusions against Linux servers remain a dominant vector of opportunistic and targeted cyber incidents, yet operational honeypots treat attacker commands either as isolated lookup keys for static templates or as input to a single uniform language model. Two research directions partially address this gap: LLM-driven response generation raises interaction realism but incurs a fidelity–latency trade-off, while semantic command analytics classifies attack stages from embeddings yet applies a single fine-tuned model uniformly across sessions. This work introduces the phase-adaptive model routing framework (PAMR), which explicitly couples both views. A lightweight online phase estimator infers the current attack stage—reconnaissance, exploitation, or persistence—from a sliding window of recent commands via compact semantic embeddings; a router then dispatches each command to one of several heterogeneous LLM backends according to estimated phase, command complexity, and a confidence-weighted latency budget. The routing decision is formulated as a constrained optimization problem and realized at runtime as an O1 lookup in a precomputed dispatch table. PAMR is evaluated on a controlled, reproducible benchmark of 412 stage-annotated command sessions that combines representative Linux command–response pairs with synthesized attacker traces; we explicitly state that this is a laboratory benchmark rather than live attacker traffic, and we scope our claims accordingly. Relative to uniform-model baselines, PAMR reduces mean response latency by approximately 38% against an API-hosted high-capacity backend and by approximately 45% against a locally hosted mid-capacity backend, while keeping token-level response-fidelity metrics close to the high-capacity baseline (cosine similarity ≈ 0.39 vs. 0.40) and maintaining an online stage-classification macro-F1 above 0.87. We further provide a first-order analytical treatment of the timing side-channel that any backend-routing architecture introduces, and we frame the contribution as a latency/token-fidelity trade-off, leaving validation of operational realism against live adversaries to future work.