Identifying Running ScientificKernels in Computing Serversthrough Execution ProfileAnalysis
Hao Bai · Uppsala University Publications (Uppsala University) · 2026
High-performance computing systems expose hardware telemetry such as RAPL-derived power readings and hardware performance counters for profiling and monitoring. These signals may also reveal work\-load-specific execution behavior. This thesis investigates whether scientific computing kernels can be identified from such execution profiles, using BLAS routines as structured test cases. The study aims to evaluate classification of both individual BLAS routines and broader BLAS levels, and to test whether learned profiles generalize to unseen problem sizes, thread counts, and routines. A benchmark suite of twelve double-precision BLAS routines was measured using RAPL power domains and hardware performance counters at a 20 ms polling interval. Run-level traces were divided into fixed windows, and statistical features were extracted from power, normalized power, IPC/cache metrics, counter features, and combined feature sets. Random Forest was used as the main classifier, with an multilayer perceptron (MLP) as a secondary baseline. Under the run-level split, exact routine classification reached 99.5% accuracy with all features, while BLAS-level classification reached up to 100%. However, performance decreased under stricter generalization settings: BLAS-level accuracy reached 94.6% for unseen size ranks, 80.6\% for unseen thread counts, and 76.7% for unseen kernels. The results show that hardware execution profiles contain discriminative information for BLAS workload identification, but robust generalization depends strongly on the type of execution variation.