Assessing and Forecasting TensorFlow Lite Model Execution Time in Different Platforms
Henrikas Giedra, Dalius Matuzevičius · 2024
Deep Neural Networks (DNNs) have emerged as powerful tools in various fields due to their ability to extract complex patterns from data. However, their computational requirements pose challenges, especially when deployed on resource-constrained devices such as mobile and embedded systems. Edge computing mitigates these challenges by employing various optimization techniques that help reduce latency. One such technique is Network Architecture Search (NAS), which automates the design of neural network architectures. This paper presents a methodology for predicting the inference time of TensorFlow Lite models, focusing on “Conv2d layers”. The TensorFlow Lite framework is chosen for its suitability for mobile and embedded devices, its minimal memory footprint, and its faster inference. Experiments conducted evaluate the performance of Convolutional Neural Network (CNN) models and predict the inference time of Conv2d layers. The results show different error margins for different configurations, providing insight into prediction accuracy and demonstrating the applicability of this methodology for predicting the execution time of different elements of CNNs on different platforms.