A Snapshot of Tiny AI: Innovations in Model Compression and Deployment
Kamal Ud-Din Cari, Marin Mercury, Guiomar Alaba · 2024
The rapid advancement of artificial intelligence (AI) has led to the development of increasingly sophisticated deep learning models capable of achieving remarkable performance across various domains. However, the deployment of these models on resource-constrained devices, such as mobile phones, embedded systems, and Internet of Things (IoT) devices, poses significant challenges due to their high computational demands and memory requirements. Tiny AI emerges as a pivotal solution to address these challenges, emphasizing the need for efficient algorithms and architectures that can operate effectively in environments with limited resources. This survey provides a comprehensive overview of the current state of Tiny AI, exploring the key techniques and methodologies employed to create lightweight AI models. We delve into model compression strategies, such as pruning, quantization, and knowledge distillation, which aim to reduce the size and complexity of neural networks while maintaining their predictive accuracy. Additionally, we examine the role of neural architecture search (NAS) in automatically designing efficient models tailored for specific applications, highlighting the trade-offs between performance and resource consumption. Furthermore, the survey addresses the importance of hardwareaware AI approaches, which optimize model deployment by considering the characteristics of target platforms. We discuss the implications of deploying Tiny AI in real-world applications, including challenges related to latency, energy consumption, and scalability. As Tiny AI continues to evolve, this survey identifies emerging trends and future directions, including the integration of advanced techniques such as federated learning and edge computing. By providing insights into the evolving landscape of Tiny AI, this paper aims to foster further research and development in creating efficient AI systems capable of operating seamlessly in diverse and resource-limited environments.