CompressionGPT: Evaluating Fault Tolerance of a Compressed Large Language Model
Neil Kapur, America Rangel, Lillian Pentecost · 2023
Deep neural networks (DNNs) currently require large amounts of memory to store weights. Consequently, inference is less efficient given that weights must be stored off-chip on DRAM, resulting in costly memory accesses. While compression techniques, including quantization and pruning, can significantly reduce model size, current memory technologies are unable to store compressed DNNs on-chip. Prior works have proposed multi-level cell emerging non-volatile memory technologies as a solution given their ability to store bits densely on-chip. While these memory technologies are fault prone, having higher bit error rates, it has been demonstrated that DNNs exhibit some fault tolerance. We build on previous work by examining the fault tolerance of a pruned and quantized large language model (LLM).