Model Optimization Techniques for Edge Devices
Yamini Nimmagadda · 2024
Chapter 4 delves into various model optimization techniques crucial for deploying AI models on edge devices such as smartphones, smartwatches, and IoT devices. These optimizations are categorized into three phases: predeployment, deployment-time, and postdeployment. Predeployment techniques include model architecture selection, quantization, structured pruning, knowledge distillation, and sparsification, which are applied to the model before production to enhance performance and efficiency. Deployment-time techniques, such as IR conversion, graph optimizations, target-dependent optimizations, dynamic batching, model caching, and model parallelism, are employed to optimize models during deployment and runtime. Postdeployment techniques, including model monitoring, retraining, hardware upgrades, and user feedback loops, ensure continuous performance improvement and adaptability of models in real-world scenarios. Through illustrative examples, this chapter provides a comprehensive understanding of how these optimization strategies can be effectively implemented to meet the constraints and requirements of edge computing.