Anomaly Classification with Unknown, Imbalanced and Few Labeled Log Data
Yulei Wu, Jingguo Ge, Tong Li · 2022
Abstract Anomaly classification plays an important role for guaranteeing the stability and reliability of IT systems. Log is an important data resource that records the state and behavior of a system. In recent years, many log-based anomaly classification methods have been proposed. However, existing methods have limitations due to the realistic nature of log data in real-world system conditions. First, most of methods rely on the complex and specific feature engineering designed by domain experts. Second, almost all supervised log analysis methods rely on the large, balanced, and labeled datasets. Third, most of methods are based on the closed-world assumption, including the “closed” of data categories, the “closed” of data scales, and the “closed” of data balance. To overcome the above limitations, in this chapter, we propose OpenLog, an anomaly classification method based on meta-learning. OpenLog uses a two-layer semantic encoder based on deep learning models to simplify the complex feature engineering. It adopts the meta-learning strategy to train the models using sufficient auxiliary datasets to enhance its performance. OpenLog transforms the multiclassification task into a binary-classification task. The target of the binary-classification task is to learn the depth relationship between samples. In this way, OpenLog can classify new anomalies without retraining due to the flexible model structure. We carry out extensive experiments using public datasets including BGL, Thunderbird, Liberty, and Spirit datasets, to demonstrate the performance of OpenLog on the unknown anomaly class data, imbalanced data, and few labeled data. Compared with several state-of-the-art methods, OpenLog shows its superiority.