GPU_HBM_initial_precooling_technical_note_english
Kawamoto Satoshi · Zenodo (CERN European Organization for Nuclear Research) · 2026
Abstract This paper presents a conceptual evaluation of predictive local pre-cooling for suppressing the initial thermal ramp around GPU-HBM regions. Conventional reactive cooling strengthens cooling after a temperature rise has already been detected, which may allow the system to enter high-temperature regions and experience throttling. Constant strong cooling can reduce temperature more aggressively, but it increases cooling energy. In this study, a short-duration local pre-cooling strategy is evaluated using leading workload features such as CPU input rate, job queue length, batch size, task type, memory intensity, compute intensity, and cache reuse factor before the temperature reaches a threshold. First, a synthetic prediction experiment was conducted to evaluate whether GPU/HBM load and thermal ramp events can be anticipated from leading features. Under simple burst conditions, CPU input rate alone detected thermal ramp events with high accuracy. However, under stress conditions including cache reuse, GPU queue delay variation, mixed jobs, burst overlap, noise, and task-type misclassification, the performance of the CPU-only model degraded. Adding job and memory-related features improved the F1 score, pre-cooling hit rate, and waste rate, suggesting that CPU-side information should be treated not as a standalone predictor but as an entry point to a broader leading feature group. Second, a simplified thermal model was used to compare reactive cooling, constant strong cooling, and initial pre- cooling. Initial pre-cooling reduced peak temperature and time above threshold compared with reactive cooling, although conventional total energy per work slightly increased. A parameter sensitivity analysis identified R1-039 as a best overall case for smoothed thermal rise rate, peak temperature, time above threshold, and energy-related metrics. A neighborhood stability analysis around R1-039 showed 70 successful cases out of 81, corresponding to a robust success rate of 86.4%. Finally, a simplified temperature-dependent throttling model was introduced to evaluate effective work. Under this throttling-aware evaluation, initial pre-cooling improved effective work done, effective work per second, average performance factor, and total energy per effective work compared with reactive cooling. It also achieved lower total energy per effective work than constant strong cooling. These results suggest that initial thermal ramp pre-cooling may contribute not only to thermal stabilization but also to effective work improvement through throttling mitigation. However, the study remains a conceptual evaluation, and validation using real GPU-HBM systems remains future work. Keywords: GPU cooling; HBM; thermal ramp; predictive pre-cooling; throttling; effective work; thermal management; simplified thermal model