Hardware Computation Graph for DNN Accelerator Design Automation Without Inter-PU Templates

Jun Li, Wei Wang, Wu-Jun Li · IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems · 2025

Existing deep neural network (DNN) accelerator design automation (ADA) methods adopt architecture templates to predetermine parts of design choices and then explore the remaining design choices beyond templates. Based on the architecture hierarchy at the processing unit (PU) level, these templates can be classified into intra-PU templates and inter-PU templates. Since templates limit the flexibility of ADA, designing effective ADA methods without templates has become an important research topic. Although there have appeared some works to enhance the flexibility of ADA by removing intra-PU templates, to the best of our knowledge no existing works have studied ADA methods without inter-PU templates. ADA with predetermined inter-PU templates is typically inefficient in terms of resource utilization, especially for DNNs with complex topology. In this paper, we propose a novel method, called hardware computation graph (HCG), for ADA without inter-PU templates. In HCG, a novel inter-PU architecture exploration strategy is proposed to optimize on-chip memory utilization. This strategy mainly depends on an appearing-frequency guided pruning method and an appearing-frequency first generation method. Experiments show that HCG can achieve competitive latency while using only 13% 90% of on-chip memory, compared with existing state-of-the-art ADA methods.

Read the paper · More papers on PaperTik