Memory Transfer Decomposition: Exploring Smart Data Movement Through Architecture-Aware Strategies
Diego A. Roa Perdomo, Rodrigo Ceccato, Rémy Neveu, Hervé Yviquel, Xiaoming Li, Jose M. Monsalve Diaz, Johannes Doerfert · 2023
Modern high-performance computing systems feature myriad compute units connected via a complex network of nonuniform links. While extensive connectivity can improve communication latency and bandwidth between components, it requires careful orchestration. However, heterogeneous programming models generally expose a flat view of the hardware where all components are connected in a star topology through a uniform, non-descriptive link to the central processing unit. This discrepancy between actual architecture and the simplified abstraction most often employed by programmers results in suboptimal utilization of the complex system interconnects.