Log Compression via Redundancy Eliminating at Word and Numerical Levels

Jiarui Chen · 2024

Logs play a crucial role in monitoring system behaviors, providing observability of system events at various levels. However, large-scale software systems generate massive amounts of logs daily, leading to tremendous storage overhead and management burdens. Compression techniques are employed to reduce the storage pressure of logs by eliminating redundancy. Traditional compression methods achieve moderate compression ratios but fail to fully exploit the structural attributes of log data. Several efforts have been made to compress logs based on their structural or temporal attributes, but they struggle with identifying word-level patterns in log contents. Addition-ally, numerical values, such as timestamps, in log headers and parameters pose challenges in compression. To address these limitations, we propose LogCompact, a novel log compressor that recognizes redundancy at both the word and numerical levels. Meanwhile, LogCompact incorporates a word encoder based on Huffman trees to reduce the encoding length of high-frequency words and a two-dimensional differential encoder to simplify the representation of numerical values. Experimental results on five datasets demonstrate the effectiveness of LogCompact in achieving higher compression rates compared to baseline methods.

Read the paper · More papers on PaperTik