eHAPPA: Efficient and Scalable Resilience Prediction in HPC Applications with Low-Rank Adaptation
Hailong Jiang, Jianfeng Zhu, Bo Fang, KEVIN J. BARKER, Chao Chen, Ruoming Jin, Qiang Guan · 2025
Predicting the soft error resilience of highperformance computing (HPC) applications is critical for deploying reliable systems, yet traditional fault injection (FI) approaches are costly and difficult to scale. We present eHAPPA, a lightweight, parameter-efficient framework that estimates resilience scores directly from source code using pretrained language models. eHAPPA builds on CodeBERT and introduces three key innovations: (1) a chunk-based encoding pipeline that handles long code sequences via tokenization, overlapping segmentation, and contextual embedding; (2) a comparative evaluation of aggregation strategies-including mean, max, LSTM, and attention pooling-to integrate chunk-level representations; and (3) the integration of Low-Rank Adaptation (LoRA), which enables scalable fine-tuning with less than 1% of parameters updated. We conduct extensive experiments across multiple configurations of LoRA rank, chunk overlap rate, and pooling method. Results show that eHAPPA consistently outperforms prior approaches such as HAPPA, achieving up to 24.6% reduction in MSE while significantly reducing computational overhead. Attention pooling combined with moderate LoRA ranks (e.g., 16) yields the best performance, demonstrating the value of expressive context modeling even under small-data regimes. Our findings suggest that pretrained models, when adapted with efficient tuning and principled aggregation, offer a powerful alternative to traditional FI-based resilience evaluation. eHAPPA provides a modular foundation for static, low-cost, and high-quality resilience estimation, with potential extensions to broader program analysis tasks such as vulnerability detection and compiler-guided optimization.