Benchmarking Small Language Models and Small Reasoning Language Models on System Log Severity Classification

Yahya Masri, Emily Ma, Zifu Wang, Joseph Rogers, Chaowei Phil Yang · Computers · 2026

System logs are essential for monitoring and diagnosing modern computing infrastructure; however, their scale and complexity require reliable automated interpretation. Although severity levels are predefined metadata, treating classification as an end task provides limited insight into a model’s ability to understand logs. Instead, severity classification can serve as a benchmark for probing runtime log comprehension. Using real-world journalctl data from Linux production servers, this study evaluates nine small language models (SLMs) and small reasoning language models (SRLMs) under zero-shot, few-shot, and retrieval-augmented generation (RAG) prompting. The results reveal clear performance stratification. Qwen3-4B achieves the highest accuracy (95.6%) with RAG, while Gemma3-1B improves substantially from 20.25% to 85.28%, and Qwen3-0.6B reaches 88.12% despite weak baseline performance. In contrast, several SRLMs exhibit performance degradation when paired with RAG. Efficiency further differentiates the models. Most Gemma and Llama variants complete inference in under 1.2 s per log, whereas Phi-4-Mini-Reasoning requires more than 228 s while achieving less than 10% accuracy. These findings indicate that architectural design, training objectives, and the ability to integrate retrieved context jointly determine performance. Overall, the results support severity classification as a practical lens for evaluating model competence and real-time deployability in digital twin systems.

Read the paper · More papers on PaperTik