A three-dimensional approach to parallel matrix multiplication

R. C. Agarwal, Susanne M. Balle, Fred G. Gustavson, Manjunath V. Joshi, Prasad Palkar · IBM Journal of Research and Development · 1995

A three-dimensional (3D) matrix multiplication algorithm for massively parallel processing systems is presented. The P processors are configured as a “virtual” processing cube with dimensions p1, p2, and p3proportional to the matrices' dimensions—M, N, and K. Each processor performs a single local matrix multiplication of size M/p1× N/p2× K/p3. Before the local computation can be carried out, each subcube must receive a single submatrix of A and B. After the single matrix multiplication has completed, K/p3submatrices of this product must be sent to their respective destination processors and then summed together with the resulting matrix C. The 3D parallel matrix multiplication approach has a factor of P1/6less communication than the 2D parallel algorithms. This algorithm has been implemented on IBM POWERparallel™ SP2™ systems (up to 216 nodes) and has yielded close to the peak performance of the machine. The algorithm has been combined with Winograd's variant of Strassen's algorithm to achieve performance which exceeds the theoretical peak of the system. (We assume the MFLOPS rate of matrix multiplication to be 2MNK.)

Read the paper · More papers on PaperTik