Optimizing Electronics Architecture for the Deployment of Convolution Neural Networks Using System-Level Modeling

Deepak Shankar, Tom Jose · 2024

As avionics transition to intelligent, image-based object detection, tracking, autonomous operations and software-defined systems, hardware architecture trade-offs for AI deployment define the cost, battery lifecycle and new safety features. The response times and power consumption of complex software such as Convolution Neural Networks vary greatly depending on the deployment at the edge vs data center, on-chip vs off-chip data storage, AI tiles vs GPU, vs CPU, partitioning into AI, software or hardware, task scheduling, and Ethernet network topology. System-level architecture exploration enables the design team to define the task graph for an operation, and model any combination of GPU, CPU, DNN, AI tiles, FPGA and custom accelerators. The user can partition individual tasks across one or more heterogeneous resources, chiplets and multiple chips. The user can optimize the mapping by running an intelligent regression that targets the requirements. Our methodology uses ChatGPT to characterize existing software and create the task graph. The table contains the tasks, data sizes, dependencies, pixel count, layers, and the sequence of execution. A set of system-level hardware IP generators of schedulers, pipelines, peripherals, and memory is used to quickly create a cycle-accurate micro-architecture hardware model. The multi-core fast simulation has an AI-diagnostic engine that detects the cause of the requirements failure and responds with alternate partitioning. We have tested this solution in the deployment of a Mask Region-Convolution Neural Network (MR-CNN) for object detection, image classification and image segmentation. We ran fifteen different configurations by varying the number of AI tiles, GPU shader cores assignment, varied the percentage of the data that is in SRAM vs external DRAM, and distributing tasks between edge vs data center. We found several interesting results including the fact that the data placement impacted performance in close to 80% of the situations. The entire project was completed in one month with simulations executed in less than 4 minutes. The accuracy was within 10% of the implemented system. The output from ChatGPT has limitations and requires manual intervention of the pixel count and number of cycles for software that does not exist. Overall, the accuracy has been sufficient for the partitioning study. ChatGPT results could be improved by training for more DNN models, both existing and proposed.

Read the paper · More papers on PaperTik