Deriving Workload Expectations - Monitoring and Analysis Using HPC Job Profiles

Joseph Fullop · OSTI OAI (U.S. Department of Energy Office of Scientific and Technical Information) · 2024

With the growing availability of time series metric data from High Performance Computing (HPC) machines, there is significant potential for using this data to improve the monitoring and analysis of HPC workloads and the systems on which they run.In this paper, we detail the statistical methods for dynamically generating job profiles and workload expectations from this data.This establishes a basis for live job monitoring and enables various methods for detecting aberrant job performance.When a deviation is detected, it can be visualized by displaying the job's metric series plotted on top of its expectation, which is represented as a cloud path.We also demonstrate how to identify system-wide anomalies by monitoring the entire set of jobs on a machine for acute, synchronized deviations.These tools can be applied to a variety of use cases such as shared resource scheduling, benchmark trending, system utilization and planning, and log analytics.In the case where jobs are not explicitly grouped by workload type, machine learning techniques can be employed to infer job types based on the shape of the time series data.Jobs are automatically classified, and when a user's workload classification changes, this indicates that something on the system (for benchmark jobs) or in the user's workload has changed.

Read the paper · More papers on PaperTik