Explainable Anomaly Detection Method Using LLM Embeddings Model with SHAP

Taisei Matsumura, Satoshi Okada, Koki Watarai, Takuho Mitsunaga · 2024

In recent years, the black-box nature of AI models in anomaly detection using log data has raised concerns about reliability and transparency. Conventional text-based anomaly detection methods have relied on structuring logs using log parsers and extracting features with techniques such as Term Frequency-Inverse Document Frequency (TF-IDF) and word2vec. How the AI model determines whether abnormal logs or not provides valuable clues for responding after detection. However, these methods depend on word frequency, making it difficult to explain the basis for anomaly detection decisions clearly. In this study, we propose a method to increase the transparency and interpretability of the model and clarify the criteria for anomaly detection by identifying features to contribute using Shapley Additive Explanations (SHAP). Our proposed method uses vectorized log data event templates using OpenAI Embeddings, which can be visually represented while achieving a high detection rate. To our knowledge, this study contributes to improving the interpretability and transparency of models by utilizing the Large Language Models (LLM) Embeddings Model for feature extraction, plotting them in a reduced-dimensional coordinate space using principal component analysis (PCA), and enhancing explainability using SHAP.

Read the paper · More papers on PaperTik