Impact of Thermal Hotspots Formation in High Performance Computing
Balvinder Pal Singh, B. Thangaraju · 2022 IEEE International Conference on Electronics, Computing and Communication Technologies (CONECCT) · 2022
In recent years computational demands have constantly increased and it is at the highest ever demand today, as almost all applications and task computations are moving into cloud-based servers. With the node size going down and the number of parallel cores on an increasing level, power density in the substrate is increasing. High power densities lead to an exponential rise in the die temperatures and high temperatures in CMOS devices degrade the performance, reliability and life of circuits. During the workload execution, local high-temperature zones (aka. Hotspots) are created on the die (in a 3D space around the material) which needs effective mitigation. Known mitigation techniques like Dynamic Thermal Management (DTM), Dynamic Voltage and Frequency Scaling (DVFS), used for managing the performance and temperature, relies on physical sensors placed onboard and are not hotspot based. Moreover, techniques are not workload aware and apply general throttling across processor cores which has a negative impact on system performance and finishing time of all jobs. Hotspot formation depends also on the physical design space of the IC, material used, and node technology used for manufacturing. This paper proposes a method to be made available to the hardware designers and software architects, at a Pre-RTL level, to model power and temperature (hotspot) based on workload and demonstrates that by applying a novel hybrid throttling policy. A hybrid scheduling policy is workload-based throttling of tasks based on the model feedback and is measured to have instantaneous power save from 20W to 2W and temperature drop from 85oC to 45oC