Self-Learning for Personalized Keyword Spotting on Ultralow-Power Audio Sensors
Manuele Rusci, Francesco Paci, Marco Fariselli, Éric Flamand, Tinne Tuytelaars · IEEE Internet of Things Journal · 2024
This article proposes a self-learning method to incrementally train (fine-tune) a personalized keyword spotting (KWS) model after the deployment on ultralow power smart audio sensors. We address the fundamental problem of the absence of labeled training data by assigning pseudo-labels to the new recorded audio frames based on a similarity score with respect to few user recordings. By experimenting with multiple KWS models with a number of parameters up to 0.5 M on two public datasets, we show an accuracy improvement of up to +19.2% and +16.0% versus the initial models pretrained on a large set of generic keywords. The labeling task is demonstrated on a sensor system composed of a low-power microphone and an energy-efficient microcontroller (MCU). By efficiently exploiting the heterogeneous processing engines of the MCU, the always-on labeling task runs in real-time with an average power cost of up to 8.2 mW. On the same platform, we estimate an energy cost for on-device training$10\times $lower than the labeling energy if sampling a new utterance every 6.1 or 18.8 s with a DS-CNN-S or a DS-CNN-M model. Our empirical result paves the way to self-adaptive personalized KWS sensors at the extreme edge.