TernaryGNNs: A High-Throughput, Area-Efficient Ternary Weight GNNs Inference Framework on CPU-FPGA Heterogeneous Platform

Zhuorong Liang, Siqi Deng, Tao Su · ACM Transactions on Reconfigurable Technology and Systems · 2025

Graph Neural Networks (GNNs) have achieved remarkable success in recent years due to their powerful ability to model non-Euclidean data structures and complex relationships. However, as graph sizes and model complexities continue to grow, the efficient deployment of GNNs across various hardware platforms presents significant challenges. The two main computation modes in GNN inference exhibit distinct characteristics in sparsity and computational density, which results in imbalanced workload distribution and inefficient utilization of compute and memory bandwidth. Furthermore, subgraph partitioning strategies, commonly adopted to address on-chip resource constraints, frequently introduce write-back overhead and port conflicts, thereby limiting overall system throughput. To address these challenges, we propose TernaryGNNs, a high-throughput, area-efficient ternary weight GNN inference framework targeting CPU-FPGA heterogeneous platforms. First, we introduce a precision-preserving ternary quantization method that maintains model accuracy with 1.016% degradation, while achieving an average weight sparsity of 81.1% and a parameter reduction rate of 94.71%. Next, by exploiting sparsity, we reformulate GNN inference as Sparse Matrix–Dense Matrix Multiplication (SpMM) or Sparse General Matrix–Matrix Multiplication (SpGEMM) computation and propose a unified sparse-optimized processor architecture. Finally, we present a comprehensive software–hardware co-design framework that ensures adaptability to the evolving landscape of diverse GNN model architectures. Our framework supports nine mainstream GNN models. Compared to the State-of-the-Art (SOTA) general-purpose GNN processor GraphOPU, TernaryGNNs achieves an average \(2.79\times\) hardware performance improvement, \(1.70\times\) end-to-end performance improvement, and \(2.83\times\) area-efficiency improvement. Compared to the leading overlay accelerator FP-GNN, it delivers \(6.73\times\) hardware performance improvement on average.

Read the paper · More papers on PaperTik