Reimagine Application Performance as a Graph: Novel Graph-Based Method for Performance Anomaly Classification in High-Performance Computing

Chase Phelps, Ankur Lahiry, Tanzima Zerin Islam, Line Catherine Pouchard · 2024

Performance anomaly in High Performance Computing (HPC) can be defined as run-to-run variation of an application in repeated runs with the same set of configuration parameters. Such variations can occur for myriad reasons, including contention for shared resources such as the network and dynamic data distribution across application processes. Traditionally HPC researchers focus on real-time anomaly detection using different Machine Learning (ML) methods. These popular methods, such as auto-encoder, limit finding anomalous event patterns during training time. On top of that, in HPC, performance data are stored in tabular format. Though gradient-based methods have already proved their significant improvement over classification tasks, they explicitly use feature-feature relationships, ignoring the potential sample-sample relationship. To fill this gap, we build a performance anomaly classification technique leveraging the potential graph-based representation learning. We hypothesize that a meaningful and robust representation considering the sample-sample relationship for the given tabular datasets will improve the downstream anomaly classification technique. We conduct our experiment on 5 HPC datasets and 6 ML datasets. Our empirical study proves that graph-based anomaly classification outperforms the gradient-based approaches in 6 out of 11 experiments. We also explain how anomaly decisions are made inside the performance graph.

Read the paper · More papers on PaperTik