Architectural support for efficient communication in scalable parallel systems

Dhabaleswar K. DK Panda, Craig B. Stunkel, Rajeev Sivaram · 1998

Efficient communication is crucial to the performance of a parallel system. Communication in a parallel system can be broadly classified into unicast (point-to-point) communication and collective communication (communication operations involving groups of nodes). To deliver good performance, a parallel system must support both forms of communication efficiently. In this thesis we propose architectural enhancements to switches to improve the performance of both unicast and collective communication. We first consider the problem of unicast communication and consider drawbacks in extant switch architectures. We then propose an input-queued switch architecture that can help overcome these problems and deliver efficient unicast communication for a range of message sizes and traffic types. We then consider two examples of frequently used collective operations: multicast/broadcast and barrier synchronization. These operations are implemented in most parallel systems on top of existent primitives for unicast communication. However, because every unicast communication incurs considerable overhead due to the software layers involved in sending and receiving a message, such an implementation is inefficient. To enhance the performance of these operations we propose a mechanism known as multidestination routing that can be used to perform efficient multicast and present the issues that must be considered for implementing such routing. We examine how two representative switch architectures can be enhanced to support deadlock-free replication and demonstrate that these schemes can improve multicast performance considerably even when compared with the best case performance of the software multicasting approaches. We then consider the problem of encoding and decoding multidestination worm headers and propose and study two schemes: multiport encoding and bit-string encoding. Next, we demonstrate how such a multicasting scheme can be used to improve multicast performance in switch-based irregular networks. We then propose architectural enhancements and protocols to make such a multicast operation reliable. Finally, we show how the mechanisms proposed for efficient and reliable multicast can be easily extended for efficient and reliable barrier synchronization.

Read the paper · More papers on PaperTik