Optimizing Communication Efficiency of GNN Inference in Distributed Systems
Wenqian Zhou, Qiaosha Zou · 2024
With the widespread adoption of Graph Neural Network (GNN) models, the demand for GNN inference acceleration has increased due to their slow runtime, especially when handling large-scale graphs. While distributed systems have shown promising results in computation acceleration, their potential for optimizing communication remains largely unexplored. In this paper, we propose novel feature pairing and feature compression methods for efficient and reduced data communication. Moreover, we co-design a multi-core system with Network-on-Chip (NoC) for evaluation. Experimental results show that our techniques can achieve up to 54.1% reduction in communication overhead, and 1.65x speedup in runtime.