CGRA4HPC 2022 Invited Speaker: Dual-scale reconfigurable arrays for ML Inference
M. Snelgrove · 2022 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW) · 2022
Untether's “Kensington” and “Boqueria” are two generations of architecture for massively parallel and highly efficient inference accelerators. They are reconfigurable at two different granularity scales: a packet scale based on precomputed circuit-switched minimum-distance Cartesian routing and a bytescale local permutation that can be reconfigured under program control. The coarse level naturally models neural-net layers and block matrices, while the permutation level carries that flexibility down to the fine structure of linear-algebra computations.