Leveraging NVML GPM for NVIDIA GPU Monitoring
Christian Wassermann, Tobias Dollenbacher, Christian Terboven, Matthias Mueller · 2026
With the continuing rise of machine learning research and applications, GPUs have become ubiquitous in the modern computing landscape. To ensure the efficient utilization of these computational resources, monitoring is often a central building block of production deployments. While nvidia-smi can provide basic utilization data, the additional metrics provided by NVIDIA’s Data Center GPU Manager (DCGM) allow more accurate insights into the actual degree of utilization. However, the integration of DCGM into an existing monitoring stack can bring its own challenges due to the limited output and deployment options. Therefore, this paper explores the GPU Performance Monitoring (GPM) metrics of the NVIDIA Management Library (NVML) that enable a low-threshold data collection of accurate and insightful GPU metrics. We validate the GPM metrics on four weeks of job data and with targeted benchmarking to clarify the interpretation of the metrics and reveal inaccuracies in the documentation. Three case studies highlight the utility of the GPM metrics for determining the application-level SM utilization and SM occupancy of GPU kernels, modeling the GPU energy consumption, and analyzing cluster workloads.