A Comparative Analysis of Loosely and Tightly Coupled Accelerator Architectures for Machine Learning
Amin Firoozshahian, Joel Coburn, Ajit Punj, Aravind Sukumaran-Rajam, Colby Boyer, Rakesh Nattoji, Mahima Bathla, Bob Dreyer, Sujith Srinivasan, Harshitha Pilla, Michael Rotzin, Surendra Rajupalem, K. Rajesh Jagannath, Krishna Noru, Harikrishna Reddy, Chris Yang, Charlie Hong-Men Su, Charlie Cheng · IEEE Micro · 2025
With the rapid innovation of AI/ML workloads, accelerators have become an important part of the computing resources in datacenters. While characteristics such as performance, power, and efficiency are critical, developer velocity, which dictates how quickly a new workload can be deployed on an accelerator, is equally important. The accelerator architecture and properties it exposes to the higher-level programming model play a key role in determining the programming model and economic viability of the accelerators. We compare two classes of AI/ML accelerator architectures: an efficient loosely coupled accelerator already implemented in silicon and a tightly coupled accelerator with a traditional and user-friendly programming model. Our comparison shows that the performance of the design depends on the amount of hardware resources and not on how they are integrated, concluding that it is possible to architect accelerators in a much simpler and more compiler-friendly manner, while still maintaining the benefits of offload computing.