Adaptive Cache Bypass and Insertion for Many-core Accelerators
Xuhao Chen, Shengzhao Wu, Li‐Wen Chang, Wei‐Sheng Huang, Carl Pearson, Zhiying Wang, Wen‐mei Hwu · 2014
Many-core accelerators, e.g. GPUs, are widely used for accelerating general-purpose compute kernels. With the SIMT execution model, GPUs can hide memory latency through massive multithreading for many regular applications. To support more applications with irregular memory access pattern, cache hierarchy is introduced to GPU architecture to capture input data sharing and mitigate the effect of irregular accesses. However, GPU caches suffer from poor efficiency due to severe contention, which makes it difficult to adopt heuristic management policies, and also limits system performance and energy-efficiency.