A reuse distance based performance analysis on GPU L1 data cache

Dongwei Wang, Weijun Xiao · 2016

Generally, cache is a bridge between CPU and main memory in order to narrow the gap of performance. As a throughput-oriented device, Graphics Processing Unit(GPU) has already integrated with cache, which is similar to CPU cores in order to exploit the locality of memory accesses. However, the applications in GPGPU computing exhibit distinct memory access patterns compared to the multi-core counterparts. Normally, the cache, in GPU cores, suffers from threads contention and resources over-utilization and few detailed works excavate the root of this phenomenon. It is significant for us to have a profound understanding of these behaviors. In this work, we adequately analyze the memory accesses from twenty benchmarks based on reuse distance theory and quantify their patterns. As a metric, the reuse distance can be employed to evaluate the cache performance and predict the access behaviors. We calculate the reuse distance for each memory access and plot them into different distributions according to the cache configuration. Through the analysis, we discover that most benchmarks either access cache in a streaming manner or reuse previous cache line in a short reuse distance. Streaming accesses barely benefit from cache since they have no data reuse. For the accesses with a short reuse distance, they can exploit the data locality in current cache design. Additionally, we discuss the optimization suggestions for all benchmarks which could improve their cache performance.

Read the paper · More papers on PaperTik