Evaluating the Data Access Efficiency of Imagine Stream Processor with Scientific Applications
Yonggang Che · 2008
The performance gap between processor and memory keeps expanding and memory access continues to be the crucial bottleneck of program performance. Traditionally, this problem is mitigated with cache technique. Stream processing is another approach that tackles this problem and has shown its effectiveness in reducing the number of memory accesses for media applications. Whether it is effective in reducing the memory traffic of scientific application is a question. This paper tries to investigate this problem. It first comparatively analyzes the memory hierarchy organization and the data access pattern of the Imagine stream processor and conventional cache based processors. Then it performs experiments on Imagine and a contrastive cache based general purpose processor (Intel Pentium M) with five typical scientific programs. The data obtained on two processors are compared against each other, with special focus on data access efficiency. The results show that data traffic between the LRF (local register file) and the SRF (stream register file) are effectively reduced on Imagine. But SRF of Imagine alone can not effectively reduce the number of off-chip memory accesses. Off-chip memory access still accounts for a large fraction of the total runtime on Imagine, as far as the programs are evaluated.