Hardware MPI Co-Processor for Reducing Communication Hotspots in Shared Memory Platforms

Ying Jun Gao, Chenlin Jin, Zhi Zheng, Letian Huang, Guozhu Liu, Hu Jun, Jinghe Wei · 2025

The bandwidth and capacity imbalance between processors and memory in traditional heterogeneous multi-core architectures is becoming increasingly pronounced. To address this issue, the industry has proposed a shared memory pool communication architecture. When implementing Message Passing Interface (MPI) point-to-point communication on such architectures, we found that directly constructing message queues in shared memory can lead to local hotspots in the on-chip network and a degradation in communication quality. Additionally, due to the lack of specific hardware support, the implementation of software mutexes introduces significant performance overhead. This paper presents an MPI Co-Processor suitable for shared memory communication. Unlike previous traditional PE-Centered MPI offloading hardware designs, we propose a Shared-Memory-Centered (SM-Centered) MPI offloading model, thereby improving scalability and resource utilization of message queues. We modified Xiong et al.'s Two-level queue and added the Wait Request Queue and Match Message Queue to support the MPI_Wait operation. By using shared-write transactions with MPI Co-Processor, PEs can avoid polling the message queue and competing for software mutexes. We conducted experiments on a multi-core heterogeneous platform. The results show that the MPI Co-Processor can significantly reduce network communication latency, effectively alleviate network hotspots, and improve communication quality.

Read the paper · More papers on PaperTik