Throughput Maximization for Transformer Inference on Processing Near-Memory Architectures

Mengke Ge, Yingjian Zhong, Song Chen, Yi Kang · IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems · 2025

The advent of Transformers has revolutionized fields such as computer vision and natural language processing. However, their memory-intensive nature creates significant hurdles for conventional computing platforms such as CPUs and GPUs. Processing near-memory (PNM) architecture has arisen as a promising solution to mitigate the memory wall problem. However, efficiently deploying Transformer models on PNM architecture remains a cutting-edge challenge. To address the practical demands of cloud and edge computing, we propose a novel mapping framework called Energon, which aims to facilitate high-throughput inference of encoderbased Transformers on PNM-based neural network (NN) accelerators, catering to both non-latency-sensitive and latencybounded scenarios. Firstly, Energon introduces a novel pipeline parallelism strategy based on an XY-aligned layout, which offers an enhanced flexibility in pipeline layout compared to existing pipeline parallelism approaches, while adapting to the finegrained partitioning scheme tailored for Transformers to achieve efficient mass parallelism. Secondly, Energon formulates the mapping optimization problems using dynamic programming and integer linear programming, respectively, to jointly optimize network partitioning and pipeline layout construction for a globally optimal mapping solution. Experimental results demonstrate that Energon significantly improves the inference throughput of encoder-based Transformers on PNM accelerators, outperforming state-of-the-art mapping frameworks by 1.1× to 2.3×. Under user-defined latency bounds, it enhances the inference throughput by an average of 43% and up to 123%.

Read the paper · More papers on PaperTik