Analysis of double buffering on two different multicore architectures: Quad-core Opteron and the Cell-BE

José Carlos Sancho, Darren J. Kerbyson · Proceedings - IEEE International Parallel and Distributed Processing Symposium · 2008

In order to take full advantage of multi-core processors careful attention must be given to the way in which each core interacts with main memory. In data-rich parallel applications multiple transfers between the main memory and local memory (cache or other) of each core will be required. It will be increasingly important to overlap these data transfers with useful computation in order to achieve highperformance. One approach to exploit this compute-transfer overlap is to use double-buffering techniques that require minimal resources in the local memory available to the cores. In this paper, we present optimized buffering techniques and evaluate them for two state-of-the-art multi-core architectures: quad-core Opteron and the Cell-BE. Experimental results show that using double buffering can substantially deliver higher performance for codes with data-parallel loop structures. Performance improvements of 1.4times and 2.2times can be achieved for the quad-core Opteron and Cell- BE respectively. Moreover, this study also provides insight into the application characteristics required for achieving improved performance when using double-buffering, and also the tuning that is required in order to achieve optimal performance.

Read the paper · More papers on PaperTik