Optimizing CNN Inference on Multicore Scratchpad Architectures

Chiara Daini, Giuseppe Lipari, Houssam-Eddine Zahaf, Pierre-Emmanuel Hladik · 2025

Many Artificial Intelligence algorithms (e.g. Convolutional Neural Networks - CNNs) can be modeled as a collection of functions which communicate with each other according to a directed acyclic graph. The main goal of this paper is to optimize the execution of CNNs on real-time embedded systems based on a multicore architecture with scratchpad memory. In a typical multicore platform, cores share a complex memory hierarchy with one or more levels of cache memories, leading to potential interference and contention on the shared communication buses. In these architectures, it is very difficult to bound the tasks' execution time and the communication delay, due to the unpredictable behavior of the cache subsystem. To reduce contention, it is possible to use architectures based on scratchpads, where every processor has a dedicated programmable fast memory to perform its local computations, and data is moved between the main memory and the local memories according to a timed schedule. In this paper, we study the problem of allocating CNN inference functions to processor cores, and scheduling the execution and memory communications. We propose an Integer Linear Programming (ILP) model that accounts for both the cost of copying data to and from scratchpad memories and the parallel computation costs on the cores, with the goal of meeting real-time temporal constraints. We propose an abstract model and two different optimization techniques, offering a trade-off between analysis time and performance. To evaluate our approach, we compare its effectiveness to classic techniques for accelerating matrix multiplication, using benchmarks from the literature. Our results demonstrate that our ILP constraints provide significant improvements in optimizing realistic CNNs.

Read the paper · More papers on PaperTik