A Task Scheduling Scheme for Preventing Temperature Hotspot on GPU Heterogeneous Cluster

Yunpeng Cao, Haifeng Wang · 2017

With the development of GPU and GPGPU, GPU heterogeneous cluster has been widely used in parallel data processing in data center. In data green and sustainable computing, power consumption is a problem worthy of consideration. Power consumption may cause temperature rise and the problems of reliability and performance will probably occur. So temperature management and controlling has become a new research hotspot. If the temperature of a node (a computer in cluster) reaches a very high value after long time running, it is called temperature hotspot. Temperature hotspot has serious influence on computing reliability and energy efficiency. In order to prevent GPU cluster temperature hotspot from occurring, a task scheduling scheme for preventing temperature hotspot was proposed. According to computing scale threshold, we use asymmetric partitioning strategy to divide computing task and schedule sub-tasks among computing nodes. In this scheme, temperature, reliability and computing performance are considered to reduce performance differences among nodes and improve throughput. By optimizing scheduling, temperature hotspots caused by slow nodes are prevented. The experimental results show that by using the proposed scheme we can control node temperature and prevent temperature hotspot while guaranteeing computing performance and reliability.

Read the paper · More papers on PaperTik