Hardware-Accelerated On-Device Learning: Training, Partitioning, and Compilation for Constrained Edge AI

Iuliia Topko, Alexey Serdyuk, Tanja Harbaum, Jürgen Becker · 2025

Real-world applications, such as autonomous driv- ing, require continuous adaptation of Deep Neural Networks (DNNs) to local environments. In addition to the growth of DNNs in recent years, Transfer Learning methods that adapt a pre-trained model to a new domain have become more widespread. A further improvement of these methods is on- device learning, which allows model adaptation based on the domain data directly on the device. However, efficient on-device training on constrained edge AI remains a challenging task, due to the limited available compute resources. Currently, the deployment process, from Machine Learning (ML) frameworks to hardware accelerators, is mainly optimized for inference. This paper proposes an end-to-end on-device learning strategy that considers three main aspects: algorithmic techniques for model adaptation, ML compilation, and hardware architectures capable of training. We present an initial model analysis highlighting potential partitioning points for hardware-accelerated on-device learning. Index Terms-Machine Learning, Transfer Learning, Ondevice Learning, Hardware Accelerator

Read the paper · More papers on PaperTik