TripleOptim: A Comprehensive Optimization Framework for GPTQ Quantization Inference on Heterogeneous Platforms

Wei Wang, Lin Han, Jiehan Zhou, Jinling Yu, Jingming Xie, Chaowei Kong, Zhenqi Gao · KSII Transactions on Internet and Information Systems · 2025

To address the performance bottlenecks in the GPTQ quantization inference on heterogeneous platforms, we conducted an in-depth analysis of the vLLM architecture, identifying linear algebra operations and fused operations as critical components for performance enhancement.Based on this observation, we proposed a comprehensive optimization framework, termed TripleOptim, designed to significantly improve both inference speed and computational efficiency.TripleOptim introduces a memory access optimization technique, termed HalfMemOpt, which leverages half-precision numbers to reduce memory access latency.It also incorporates HighVecOpt, a high-precision vectorized floating-point arithmetic optimization, to improve computational efficiency through parallelism.Lastly, TripleOptim applies an instruction-level optimization strategy, InstrMatOpt, to accelerate linear algebra operations, further streamlining the inference process.Experimental results show that TripleOptim significantly boosts throughput across all models, achieving increases ranging from 34% to 98%.These optimization strategies not only accelerate model inference speed but also deliver substantial enhancements in inference performance on heterogeneous platforms.

Read the paper · More papers on PaperTik