Speaking style compensation on synthetic audio for robust keyword spotting
Houjun Huang, BYanmin Qian · 2022 13th International Symposium on Chinese Spoken Language Processing (ISCSLP) · 2022
With the rise of intelligent speech processing applications, quickly producing keyword spotting (KWS) models with low resource has gained particular importance in recent years. Multi-speaker text-to-speech (TTS) has been proved to be an effective data augmentation technique for KWS to help complement inadequacies in the training data. However, in previous works, KWS system built with TTS augmented data couldn’t obtain considerable performance of that trained with real recordings as synthetic speeches could not fully represent target speaker’s speaking style. This work focuses on how to make synthesized speeches to be more similar to reference speaker’s speaking style under a specific metric. Speaker classification accuracy of synthesized keyword data on a speaker recognition model trained with real common recordings is used as the objective metrics, and the deep complex convolution recurrent network (DCCRN) is used to optimize it. Experimental results show that TTS augmentation helps improve the KWS system’s robustness. Moreover, by compensating speaking style of the synthetic data, we achieve a significant further improvement.