Computation and Complexity-Aware DNN Accelerators: Architectures, Dataflows, and Design Trade-Offs
Sudheer Vishwakarma, Zaheer Khan, Janne J. Lehtomäki · IEEE Access · 2026
The rapid advancement of deep neural networks (DNNs), driven by increasingly complex models and demanding computational workloads, has created a critical need for specialized hardware accelerators that effectively balance performance, energy efficiency, and scalability. This paper presents a comprehensive survey of DNN accelerator architectures, focusing on design principles that are both computationally intensive and aware of complexity.We systematically integrate various accelerator designs, including platform-based architectures, application-specific integrated circuits (ASICs), and data flow-driven accelerators, along with emerging paradigms. The importance of these designs is highlighted in relation to the challenges of computational intensity and model complexity. Additionally, we introduce an overview that encompasses key architectural dimensions, including the organization of processing elements (PEs), memory hierarchy, granularity of parallelism, and data flow strategies such as weight-stationary, output-stationary, row-stationary, and systolic dataflows. The paper provides an in-depth analysis of both seminal and contemporary ASIC-based accelerators (DianNao family, Eyeriss, Cambricon-X), data flow-driven accelerators, and Edge vs. cloud accelerators.