Towards efficient deep learning for vision and language applications
Zhenglun Kong · 2024
Deep neural networks (DNNs) have become a fundamental element and core enabler of artificial intelligence. More recently, the emergence of large language and vision models has significantly advanced the AI frontier. This progression has led to many incredible, cutting-edge applications across various fields, enhancing daily life in areas such as community/shared virtual reality experiences, self-driving cars, healthcare and medical research, accessibility and assistive technologies, education, business, and more.Much of the advancements crucially depend on the fast execution of high-quality deep learning models. In the diverse landscape of AI-associated platforms, mobile, and embedded computing devices have emerged as pivotal carriers of deep learning. These devices are extensively wearable devices, robotic vision and control, and smart health devices, thereby facilitating the spread of machine intelligence. In this dissertation, the focus is on achieving the general and practical implementation of AI, primarily through the development of efficient and robust AI models and algorithms for various applications. The high computation and storage demands of executing DNN training and inference are critical challenges that may impede the deployment of AI applications. For example, large transformer-based models, which consist of self-attention layers capable of capturing long-range dependencies and complex patterns in data, present significant challenges for efficient training and inference. These challenges include high computational cost, large memory footprint, and low hardware utilization. This dissertation will address how these key AI challenges are tackled from four main directions: managing massive computation, reducing training costs, designing speed-aware efficient models for practical deployment, and model integration for enhanced performance and efficiency. --Author's abstract