Shared Receive Queue Based Scalable MPI Design for InfiniBand Clusters

Sayantan Sur, Lei Chai, Hyun‐Wook Jin, Dhabaleswar K. DK Panda, Sun Microsystems · 2006

Clusters of several thousand nodes interconnected with InfiniBand, an emerging high-performance intercon-nect, have already appeared in the Top 500 list. The next-generation InfiniBand clusters are expected to be even larger with tens-of-thousands of nodes. A high-performance scalable MPI design is crucial for MPI appli-cations in order to exploit the massive potential for paral-lelism in these very large clusters. MVAPICH is a popular implementation of MPI over InfiniBand based on its reli-able connection oriented model. The requirement of this model to make communication buffers available for each connection imposes a memory scalability problem. In or-der to mitigate this issue, the latest InfiniBand standard in-cludes a new feature called Shared Receive Queue (SRQ) which allows sharing of communication buffers across mul-tiple connections. In this paper, we propose a novel MPI de-sign which efficiently utilizes SRQs and provides very good performance. Our analytical model reveals that our pro-posed designs will take only 1/10th the memory require-ment as compared to the original design on a cluster sized at 16,000 nodes. Performance evaluation of our design on our 8-node cluster shows that our new design was able to provide the same performance as the existing design while requiring much lesser memory. In comparison to tuned ex-isting designs our design showed a 20 % and 5 % improve-ment in execution time of NAS Benchmarks (Class A) LU and SP, respectively. The High Performance Linpack was able to execute a much larger problem size using our new design, whereas the existing design ran out of memory.

Read the paper · More papers on PaperTik