Overview Of Tensor Layout In Modern Neural Network Accelerator
Wang Zhongyu Tao, Yuanfeng Wang, Huaisheng Zhang · 2021 18th International Computer Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP) · 2021
Modern neural network accelerator is widely used to accelerate the training or inference in neural network related works, such as CNN, RNN, NLP, etc. The tensor is the fundamental unit to store the neural network's data. In the paper, we profile two prevalent tensor layout NHWC and N(C/x)HWx to see different layout's influence to L1/L2 cache. We design a profiling framework which includes block splitting engine, tensor data loading engine and memory with hierarchy structure. Experimental results can show the different L1/L2 cache configurations' influence for cache hit/miss count, load/store count and evict count, etc., which can guide the basic hardware design for L1/L2 cache in modern neural network accelerator.