Design of Chiplet-based Domain-Specific Architecture Accelerator with Scalable Mac Array
Dongqing Fang, Zhensong Li, Bingjie Li, Wenbo Mu, Chenghao Yuan, Liang Zhang · 2025
As the latest carrier for More than Moore, chiplet technology has become an important technology to break through development bottlenecks in recent years. Meanwhile, the design of Domain-Specific Architecture (DSA) accelerator has been a popular research focus. In order to implement a high-speed inference of VGG16-int8 model for real-time tasks in edge device, a network-on-chip(NoC)-interconnected chiplet-based domain specific accelerator architecture is explored in this paper. This architecture eliminates not only a large number of memory access processes in data computation and transmission, but also more crucially, improves computational efficiency, and optimizes power consumption. Simultaneously, a practical chiplet of a Mac array Core (MC) is designed on SMIC 180 nm CMOS process node to realize the most significant convolution module of VGG16-int8 model. The implementation involved synthesizing the Verilog code of the MC module, completing layout placement, routing, physical design rule checks (DRC), and layout versus schematic (LVS) verification. The finalized design layout was submitted to the multi-project wafer (MPW) host institution for mask stitching, followed by the fabrication of prototype chips for basic functional validation. This process successfully established the entire DSA accelerator design, workflow, achieving full-cycle integration from concept to realization. Simulation results show that the interconnection chiplet architecture designed in this paper can meet the needs of deep learning acceleration and provide an effective solution for chiplet integration.