Opal: A 16nm Coarse-Grained Reconfigurable Array for Full Sparse ML Applications
Po‐Han Chen, Bo Wun Cheng, Michael Oduoza, Zhouhua Xie, Kalhan Koul, Sai Gautham Ravipati, Yuchen Mei, Rupert Lu, Alex Carsello, Mark Horowitz, Priyanka Raina · 2025
Machine learning (ML) models have been rapidly growing in size to improve accuracy, but this has greatly increased their computational cost. Leveraging the unstructured sparsity in the models' inputs and weights is a promising way to reduce computational cost. However, prior sparse accelerators [1], [2] have targeted specific sparse models, and as ML models evolve, these accelerators quickly become outdated. Coarse-grained reconfigurable arrays (CGRAs) can adapt to the fast-changing nature of ML models with their reconfigurable data path. However, prior CGRAs [3], [4] face several performance challenges when they execute sparse ML models, as shown in Fig. 1. Firstly, prior CGRAs support executing sparse matrix multiplication (spMspM) using only the inner product dataflow. Gustavson's spMspM algorithm [5] provides improved performance in sparse computations by skipping rows that do not contribute to the result; however, it does not run on prior CGRAs due to its dependence on a vector reduction dataflow, which is not supported. This limitation results in a missed opportunity for performance improvement, especially considering that spMspM is a critical kernel in sparse ML applications. Secondly, the microarchitecture of prior CGRAs was designed to process long data streams. However, this assumption does not hold for sparse applications that often operate on short data bursts, making prior work vulnerable to pipeline stalls. Thirdly, prior accelerators experience significant runtime overhead due to the lack of native support for complex operations such as softmax and dense-to-sparse conversion. This results in higher data movement, impacting the end-to-end ML application performance.