Automated Partitioning of Data-Parallel Kernels using Polyhedral Compilation

Alexander Matz, Johannes Doerfert, Holger Fröning · 2020

GPUs are well-established in domains outside of computer graphics, including scientific computing, artificial intelligence, data warehousing, and other computationally intensive areas. Their execution model is based on a thread hierarchy and suggests that GPU workloads can generally be safely partitioned along the boundaries of thread blocks. However, the most efficient partitioning strategy is highly dependent on the application’s memory access patterns, and usually a tedious task for programmers in terms of decision and implementation.

Read the paper · More papers on PaperTik