Implementing High-Performance Complex Matrix Multiplication via the 1M Method
Field G. Zee · SIAM Journal on Scientific Computing · 2020
Almost all efforts to optimize high-performance matrix-matrix multiplication have been focused on the case where matrices contain real elements. The community's collective assumption appears to have been that the techniques and methods developed for the real domain carry over directly to the complex domain. As a result, implementors have mostly overlooked a class of methods that compute complex matrix multiplication using only real matrix products. This is the second in a series of articles that investigate these so-called induced methods. In the previous article, we found that algorithms based on the more generally applicable of the two methods---the 4m method---lead to implementations that, for various reasons, often underperform their real domain counterparts. To overcome these limitations, we derive a superior 1m method for expressing complex matrix multiplication, one which addresses virtually all of the shortcomings inherent in 4m. Implementations are developed within the BLIS framework, and testing on microarchitectures by three vendors confirms that the 1m method yields performance that is generally competitive with solutions based on conventionally implemented complex kernels, sometimes even outperforming vendor libraries.