SuperHCA: An Efficient Deep-Learning Edge Super-Resolution Accelerator With Sparsity-Aware Heterogeneous Core Architecture
Zhicheng Hu, Jiahao Zeng, Xin Zhao, Liang Zhou, Liang Juan Chang · IEEE Transactions on Circuits and Systems I Regular Papers · 2024
Deep learning-based super-resolution (SR) generative models have recently emerged as a promising approach for generating high-quality images. While large SR networks can achieve a high peak signal-to-noise ratio (PSNR) to assess image quality, they often come with a high number of parameters, leading to increased computational and memory requirements that can be challenging to deploy on embedded hardware. In this study, we introduce the Anchor-Based Shuffle Net (ABSN), which is designed to create a hardware accelerator using a dynamic-scale fixed-point (DSFP) quantization method. Additionally, we incorporate dynamic quantization adaptation in the hardware design. Our Super-resolution Heterogeneous Accelerator, SuperHCA, utilizes a sparsity-aware heterogeneous architecture to optimize inference efficiency by distinguishing between dense and sparse workloads. We also propose Slice Layer Fusion (SLF) dataflow and feature-sharing bit interleaving (FSBI) methods in the heterogeneous cores to reduce on-chip buffer sizes. The SuperHCA achieves a frame rate of 91 fps at a target resolution of FHD, with the highest throughput area ratio (TAR) of 22.75 fps/mm2 compared to existing state-of-the-art works.