Power-Aware, Value-Opitimized LSF Scheduler for CFD Jobs on Large HPC Clusters
Roy Campbell · 44th AIAA Aerospace Sciences Meeting and Exhibit · 2006
*As the number of processors in a typical high performance computing (HPC) system grows, practical yet subtle issues arise, some of which may possibly be mitigated by imposing a feedback control loop on a given system’s resident job scheduler. The first issue stems from an increased complexity in the underlying infrastructure that supplies power to the HPC system. The amount of power delivered to today’s large systems often warrants multiple uninterruptible power supplies (UPSs). Therefore, maximizing the span time of battery backup (in the event of an outage) requires a strategic allocation of work across the target system such that the demand for backup power is balanced across the divided power infrastructure. The second issue concerns execution value. As scheduling spaces grow larger, automated methods will become increasingly necessary to ensure that queuing attributes are dynamically tuned to maximize overall value. This paper discusses the major issues faced in making a scheduler both power and value aware. The actual construction of such a scheduler (tied to LSF) is reserved as a future work. To aid in the discussion, the effect of imposing the DoD HPC Modernization Program’s value metric on a standard queuing system (of a large cluster) is explored via simulation.