Computational Complexity of Neural Network Linear Layer Inference Optimizations

Klaus Pendl, Branislav Rudić · 2024

The deployment of machine learning models on devices with limited memory and processing capabilities demands interdisciplinary expertise in the areas of data science, algorithms and computer architecture. We present a thorough analysis for inference optimization in linear layers, which are fundamental in many models and can induce computational bottlenecks. The considerations are detached from the executing hardware, which allows for a detailed understanding of the computations involved on an arithmetic level. We conclude with appropriate metrics and expressions, which are derived based on concepts of computational complexity, and provide quantitative or implicit measures of the linear layer's performance and resource usage. The framework is systematically applied to various optimization methods, each offering a different approach to reduce the computational complexity of linear layers and thus the overall model, and can therefore be crucial for deployment on resource limited devices.

Read the paper · More papers on PaperTik