WOLF: Weight-Level OutLier and Fault Integration for Reliable LLM Deployment
Chong Wang, Wanyi Fu, Jiangwei Zhang, Shiyao Li, Rui Hou, Jian Yang, Yu Wang · IEEE Transactions on Computers · 2025
The rapid advancement of Transformer-based large language models (LLMs) is presenting significant challenges for their deployment, primarily due to their enormous parameter sizes and intermediate results, which create a bottleneck in memory capacity for effective inference. Compared to traditional DRAM, Non-Volatile Memory (NVM) technologies such as Resistive Random-Access Memory (RRAM) and Phase-Change Memory (PCM) offer higher integration density, making them promising alternatives. However, before NVM can be widely adopted, its reliability issues, particularly manufacturing defects and endurance faults, must be addressed. In response to the limited memory capacity and reliability challenges of deploying LLMs in NVM, we introduce a novel low-overhead weight-level map, namedWolf.Wolfnot only integrates the addresses of faulty weights to support efficient fault tolerance but also includes the addresses of outlier weights in LLMs. This allows for tensor-wise segmented quantization of both outliers and regular weights, enabling lower-bitwidth quantization. TheWolfframework uses a Bloom Filter-based map to efficiently manage outliers and faults. By employing shared hashes for outliers and faults and specific hashes for faults,Wolfsignificantly reduces the area overhead. Building onWolf, we propose a novel fault tolerance method that resolves the observed issue of clustering critical incorrect outliers and fully leverages the inherent resilience of LLMs to improve fault tolerance capabilities. As a result,Wolfachieves segment-wise INT4 quantization with enhanced accuracy. Moreover,Wolfcan adeptly handle Bit Error Rates as high as$1 {\boldsymbol{\times}} 10^{-2}$without compromising accuracy, in stark contrast to the state-of-the-art approach where accuracy declines by more than 20%.