Dynamic Adapting Scheduling for HPC: Eliminating Job Failure through Robust Resource Allocation

Vishesh Goyal, N. Pavithra · 2025

This paper presents a dynamic and intelligent workload scheduling framework tailored for high-performance computing (HPC) environments. Unlike traditional models that rely on static assumptions or failure prediction, our approach leverages real-time system monitoring, dependency-aware prioritization, and adaptive scheduling to improve fault tolerance and execution efficiency. We introduce a multi-layered strategy featuring calculated job priorities, resource availability checks, dependency-based job promotion, limited parallel execution, and a starvation-preventing waiting queue mechanism. Evaluated on a synthetic dataset of 5000 jobs, the proposed scheduler achieved superior results compared to classical approaches such as First-Come-First-Served (FCFS), Shortest Job First (SJF), and predictive failure models. Our model reduced makespan and waiting times while maintaining lower dependency violations, all without requiring historical data for failure prediction. These results demonstrate the viability of our method as a scalable and adaptive alternative for managing complex HPC workloads.

Read the paper · More papers on PaperTik