Robust Representations for Keyword Spotting Systems

Aidan Smyth, Niall Lyons, Ted S. Wada, Robert Zopf, Ashutosh Pandey, Avik Santra · 2022 26th International Conference on Pattern Recognition (ICPR) · 2022

Keyword spotting poses several challenges due to acoustic disturbances such as noise, reverberation, and speaker-to-speaker variations. Typical keyword spotting features such as Log-Mel spectrograms or Mel Frequency Cepstral Coefficients are highly sensitive to frequency perturbations. Denoising such features using an autoencoder improves performance in noise but often results in poor generalization for unseen speakers. This paper proposes an architecture that addresses the above challenges for keyword spotting systems. The audio features are combined with a latent representation extracted by a denoising autoencoder. In addition, this paper proposes a novel method for creating highly separable optimal latent representations from speech using a discriminative denoising autoencoder trained with a quadruplet loss metric learning approach. The proposed approach creates a discriminative latent representation, which when combined with the original input results in an architecture that is ideal for keyword spotting. The proposed architecture outperforms all approaches when tested in both clean and noisy environments with reverberation at various distances on unseen speakers.

Read the paper · More papers on PaperTik