HPC Monitoring & Visualization

Jaelyn Litzinger, Roy Hallquist, James Tessmer · Practice and Experience in Advanced Research Computing · 2023

How do you check if you're using the HPC resources you've allocated? Are your multi-GPU jobs using all allocated GPUs or is the job just overloading the first GPU while the others remain idle? To help users answer these questions for themselves, we have engineered a monitoring service for our main HPC cluster's CPU and GPU resources. This service includes a live dashboard that graphs time-series metrics like memory usage, disk read/write times, network traffic, etc. These metrics can also be queried for their raw numerical data through a Python script for further analysis.

Read the paper · More papers on PaperTik