Exploring Constrained Dataflow Accelerators for Real-Time Multi-Task Multi-Model Ml Workloads
Jamin Seo, Jianming Tong, Tushar Krishna, Hyoukjun Kwon · 2025
Emerging machine learning (ML) workloads, such as those in AR/VR applications or drones, exhibit real-time multitask multi-model (RT-MTMM) characteristics. These workloads demand efficient executions of diverse combinations of multiple models within strict real-time constraints and tight energy budgets. Consequently, optimizing hardware accelerator design and dataflow choices for such ML models becomes imperative. Flexible dataflows in hardware have been explored to enhance performance and energy efficiency for diverse models. However, this flexibility involves substantial hardware costs, potentially outweighing the performance benefits. While coarse-grained flexible accelerators, including heterogeneous dataflow accelerators [1], have been proposed to mitigate such costs, they may prove suboptimal depending on combinations of models, potentially leading to under-performance of the hardware. Hence, this paper investigates the desired balance between the benefits and costs of flexibility. We present a systematic and quantitative exploration of medium-grained flexible accelerators, demonstrating that constrained yet judicious domain-aware flexibility choices in tile sizes (T), loop order (O), parallel dimensions$(\mathrm{P})$, and array shape (S) can allow RT-MTMM accelerators to achieve comparable real-time performance ($\times 1.06 / \times 1.13$) and energy efficiency ($\times 1.09 / \times 1.07$) to fully flexible designs with significantly lower area overhead ($\times 0.67 \times 0.91$) on edge/mobilescale accelerators.