SGRACE: Scalable Architecture for On-Device Inference and Training of Graph Attention and Convolutional Networks

Jose Luis Nunez-Yanez, Hadi Mousanejad Jeddi · IEEE Transactions on Very Large Scale Integration (VLSI) Systems · 2025

In this article, we propose a hardware accelerator for on-device inference and training of deeply quantized graph convolutional networks (GCNs) and graph attention networks (GAT). The architecture unifies in a single dataflow both GAT and GCN support without breaking the matrix global formulation needed for high performance. In addition, it supports adaptive fixed-point precision (from 1- to 8-bit) and provides scalable performance through configurable hardware threads and compute units available in each thread. Each hardware thread employs a hierarchical dataflow combining fine- and coarse-grained components. The fine-grained dataflow streams words with a bit-width that depends on the selected precision. The coarse-grained dataflow connects aggregation and combination stages, incorporating an optional attention mechanism in GAT mode which is bypassed in GCN mode. In training mode, the hardware emulates multiple precisions (i.e., 1-/2-/4-/8-bit) to perform hardware-aware quantized training (HQT) for weights, features, attention coefficients, and the graph connectivity expressed in the normalized adjacency matrix. In inference mode, the architecture is optimized to support a single precision (i.e., 1- and 2-bit). The accelerator is mapped to a Zynq Ultrascale MPSOC and integrated within the Pytorch framework. Performance evaluation on Planetoid and Molecular graph datasets demonstrates over$100\times $acceleration for GCN and GAT compared to on-device CPU execution. The comparison with other FPGA, mobile GPU, and CPU hardware shows up to$50\times $better performance per joule. Results confirm HQT as an effective and efficient quantization strategy for graph neural network acceleration.

Read the paper · More papers on PaperTik