Asymmetric NoC Architectures for GPU Systems
Amir Kavyan Ziabari, José Luis Abellán, Yenai Ma, Ajay M. Joshi, David R Kaeli · 2015
While both Chip MultiProcessors (CMPs) and Graphics Processing Units (GPUs) are many-core systems, they exhibit different memory access patterns. CMPs execute threads in parallel, where threads communicate and synchronize through the memory hierarchy (without any coalescing). GPUs on the other hand execute a large number of independent thread blocks and their accesses to memory are frequent and coalesced, resulting in a completely different access pattern.