On-chip Memory in Accelerator-based Systems: A System Technology Co-Optimization (STCO) Perspective for Emerging Device Technologies
Siva Satyendra Sahoo, Dawit Burusie Abdi, Julien Ryckaert, James Myers, Dwaipayan Biswas · 2024
The last few years have witnessed exponential growth in the deployment of AI/ML-based processing across application domains. As a result, domain-specific accelerators for AI are being implemented across different scales of computing—from edge to cloud computing. As the complexity of emerging AI/ML models increases exponentially, off-chip data access remains a major bottleneck for AI accelerators. Increasing On-chip Memory (OCM) capacity, to improve data reuse remains one of the primary approaches to reducing costly off-chip DRAM data access. Similarly, devices enabling dense memory arrays are being actively explored to scale on-chip memory capacity within limited resource constraints. However, most related works perform such an analysis without exploring the impact of the AI model’s complexity on the resulting power, performance and area (PPA) trade-offs. To this end, we present a System-technology co-optimization (STCO) study of the impact of OCM scaling on modern AI workloads. Specifically, we present a joint analysis across workload complexity, systolic array accelerator scaling, and OCM device technology options.