NUPEA: Optimizing Critical Loads on Spatial Dataflow Architectures via Non-Uniform Processing-Element Access

Souradip Ghosh, Graham Gobieski, Keyi Zhang, Brandon Lucia, Nathan Beckmann, Tony Nowatzki · 2025

Data movement is the dominant energy, performance, and scalability bottleneck in modern architectures.Systems have tackled data movement by distributing data, e.g., via non-uniform memory access (NUMA) architectures.However, to reduce data movement, these architectures must identify critical data and place it closer to compute.Clever data placement is complex and often ineffective.Spatial dataflow architectures (SDAs) present a new opportunity to tackle data movement.SDAs distribute program instructions across a spatial fabric of processing elements (PEs).On large SDAs, some PEs are necessarily closer to memory than others, giving rise to non-uniform processing-element access (NUPEA).Clever instruction placement can thus reduce data movement by, e.g., placing critical loads close to memory.This paper introduces NUPEA and contrasts it with prior datacentric approaches to scaling data movement.We find that it is often easier for the compiler to identify critical loads than the data they access, making NUPEA applicable where NUMA is not.We present simple architecture and compiler optimizations for NUPEA and implement them on the Monaco SDA architecture and effcc compiler, both industry products by Efficient Computer.On Monaco, across a range of important kernels, NUPEA yields an avg 28% speedup over a uniform-PE-access (UPEA) SDA and an avg 20% speed over a UPEA SDA with NUMA.

Read the paper · More papers on PaperTik