Cluster Heartbeat: A Unified Behavioral Fingerprint for GPU Cluster Reliability, Scheduling, and Cost Optimization
Nahid Ibn Zaman · Zenodo (CERN European Organization for Nuclear Research) · 2026
Modern GPU clusters are typically monitored, scheduled, and cost-optimized by three independent systems, each observing only a partial slice of node behavior. We present Cluster Heartbeat, a system that compresses multi-metric GPU telemetry (13 features spanning utilization, thermal, power, error-rate, and I/O signals) into a single learned workload fingerprint using a multi-head PyTorch autoencoder. One fingerprint feeds three downstream services — predictive failure detection, behavior-aware scheduling, and idle/ghost-job cost optimization — from the same shared representation. On a synthetic but physics-plausible DCGM telemetry benchmark (16 nodes, injected incidents), the system achieves 0.78 AUROC and 0.67 F1 (0.96 precision) for anomaly detection, a mean predictive lead time of approximately 36 minutes before node failure (up to 1.1 hours for ECC error bursts), 97% workload classification accuracy, and 6.6 percentage-point MAE on next-step GPU utilization prediction.