A Novel Approach to the Detection of Stegomalware in PDFs Using Kolmogorov Complexity

Sebastian Alexis · 2025

PDF-based stegomalware is a growing cybersecurity threat that disguises malicious payloads within the tree structure of PDF documents, exploiting features like embedded JavaScript, metadata, and open actions while maintaining the outward appearance of a legitimate file. Traditional detection methods (both heuristic and signature-based), often fail to identify stegomalware due to the external similarities between clean and infected documents.This project presents a novel method for detecting PDF-based stegomalware using Kolmogorov complexity. Kolmogorov complexity measures the shortest possible program to describe an output, but it cannot be computed directly. However, this project indirectly calculated complexity by analyzing the compressed lengths of attributes using compression algorithms that group repetitive (non-complex) patterns. By calculating the complexity of components within a PDF’s internal tree structure and comparing them to baselines derived from known clean documents, malicious modifications can be identified by detecting deviations from expected values.For testing, a local mail server was developed to process incoming PDFs by calculating and comparing Kolmogorov complexity to stored baselines of clean PDFs tagged with matching metadata IDs. PDFs with significant complexity increases within suspicious attributes are quarantined using Firejail, while others are safely delivered to the user.Initial testing achieved a 100% true positive rate but an 86.2% false positive rate when analyzing the entire PDF as a whole. Refining the approach to analyze specific PDF components, focusing on suspicious changes, and ignoring innocent variations like file size or text content improved results to a 97.8% true positive rate and a 3.7% false positive rate.

Read the paper · More papers on PaperTik