A Distributed Storage System for System Logs Based on Hybrid Compression Scheme
Baoming Chang, Fengxi Zhou, Zhaoyang Wang, Yu Wen, Boyang Zhang · 2023
Modern enterprises are facing a massive threat from Advanced Persistent Threats (APTs), which have risen to be one of the most dangerous challenges in recent years. Since system logs capture the complex causality dependencies between system entities, they have become the primary data source for countering APTs. However, as modern computer systems get more complicated, system logs can pile up in large quantities. Besides, APTs are sophisticated and persistent cyber attacks that can remain hidden in the target for a long time and constantly steal private data. System logs need to be collected and stored for a long duration to enable a complete analysis of APTs. Such a vast amount of log data is challenging for enterprises to store and manage. There are two mainstream solutions for reducing storage overhead. Data compression methods provide an intuitive idea. However, they are designed for general text and lack optimization for system logs. Another solution considers log reduction, which removes redundant system events recorded in system logs by predefined rules. Unfortunately, they are tailored for specific kinds of redundant information, resulting in limited applicability. Realizing that these two solutions adopt two distinct perspectives to reduce storage overhead, they are complementary. Data compression methods shrink the size of log data from their binary form. Log reduction starts from the semantic information of system logs and removes redundant information to reduce storage overhead. Combining both methods maximizes storage efficiency. In this paper, we propose a distributed storage system based on a hybrid compression scheme. To address the above deficiencies, we first identify and merge redundant system events by analyzing and tracing the information flow rather than based on rules. Then, we apply log parsing to preprocess log entries for further storage efficiency. Besides, we design a distributed architecture to optimize compression and eliminate repeated data processing steps. Our system is evaluated on a large public dataset, the results show that our system filters out 37.39% of redundant system events on average. And our hybrid compression scheme achieves the highest $56.41 \times$ and the lowest $40.31 \times$ compression ratio. In conclusion, our system achieves much better space savings compared to existing works.