Microcontroller Based Intelligent Chinese-Speech Keywords Detector by Transferring the Mid-level Features of Deep Speech
Junaid Hussain Muzamal, Muhammad Zubair Asghar, Anne Kwong, Usman Ahmed Raza · 2021
Speech Keywords Detection (SKD) can be described as a task of finding keywords in audio streams, while only a few samples of the keyword are available. SKD has immense applicability in terms of making the temporal signature of audio, to implement the word clouds, voice-command-based appliances, and intelligent voice-based agents. Existing SKD applications suffer from high power requirement constraints, false-positive results due to low training data, and missing keywords in languages like Chinese. Besides, SKD is an enormously imbalanced classification delinquent. In this work, we proposed to implement a lower power consumer and low memory occupying solution, which is built on microcontroller-based applications. We addressed the false-positive and data imbalancing problem by employing a novel combination of objective functions metric loss and Prototypical loss with the mid-level feature engineering of Deep Speech. The extensive experiments on real-time microcontroller showed that our technique is state of the art and provide comparable results with 11% improvement in F1 score.