Hierarchical Self-Supervised Learning for Semantic Embeddings of Crash Dumps

Danny Tang, Ludvig Lindholm · Lund University Publications Student Papers (Lund University) · 2026

Analyzing software crash dumps is important for diagnosing system failures, but manually grouping them by root cause is time-consuming and error-prone.At Schneider Electric, crash dumps are currently identified using file content hashing, which leaves most crashes unclassified, as any change yields a distinct crash ID.This thesis proposes a machine-learning-based embedding system that maps an entire crash dump into a single semantic embedding.Ideally, this embedding system automatically groups all crash dumps originating from the same underlying defect by assigning them similar embeddings.Since only a small fraction of the crash dumps are labeled with underlying defects, we opted for primarily using self-supervised training to maximize the amount of available data to train on.Instead of directly training the model to connect a crash dump to a defect, we train it to position crash dumps within a learned embedding space.To handle large crash dumps, a hierarchical encoder-based embedding model is introduced.Crash dumps are processed in smaller chunks, which are then combined into a single vector using learned, adaptive weighting.A novel attentionbased loss function is introduced to improve training using the learned adaptive weights, together with a multi-phase training curriculum to reduce overfitting of the attention weights.Our purely self-supervised model significantly outperforms a baseline pretrained ModernBERT model, nearly halving the False Positive Rate at a 95% retrieval threshold, which shows that the model is feasible to use as a retrieval tool for clustering.Due to extreme ground truth data scarcity, our supervised model is trained on the entire testing dataset and results are thus not directly comparable.However, the supervised model outperforms the self-supervised model on manual retrieval testing in real-world use cases.

Read the paper · More papers on PaperTik