Memory characteristics of iterative methods

Christian H. Weiß, Wolfgang Karl, Markus Kowarschik, Ulrich Rüde · 1999

Conventional implementations of iterative numerical algorithms, especially multigrid methods, merely reach a disappointing small percentage of the theoretically available CPU performance when applied to representative large problems.One of the most important reasons for this phenomenon is that the current DRAM technology cannot provide the data fast enough to keep the CPU busy.Although the fundamentals of cache optimizations are quite simple, current compilers cannot optimize even elementary iterative schemes.In this paper, we analyze the memory and cache behavior of iterative methods with extensive profiling and describe program transformation techniques to improve the cache performance of two-and three-dimensional multigrid algorithms.This project is partially funded by DFG Ru 422/7-1,2.1 All benchmarks in the article were compiled with native FORTRAN77 compilers and aggressive optimizations enabled.On the Intel platform we used egcs (V2.91.60).The platforms include an Intel PentiumII Xeon PC (450 MHz, 450 MFLOPS), a SUN Ultra 60 (296 MHz, 592 MFLOPS), a HP SPP2200 Convex Exemplar Node (200 MHz, 800 MFLOPS), a Compaq PWS 500au (500 MHz, 1 GFLOPS), and a Compaq XP1000 (500 MHz, 1 GFLOPS).1

Read the paper · More papers on PaperTik