Accelerating DNNs inference with predictive layer fusion
MohammadHossein Olyaiy, Christopher Ng, Mieszko Lis · 2021
Many modern convolutional neural neworks (CNNs) rely on bottleneck block structures where the activation tensor is mapped between higher dimensions using an intermediate low dimension, and convolved with depthwise feature filters rather than multi-channel filters. Because most of the computation lies in computing the large dimensional tensors, however, such networks cannot be scaled without significant computation costs.