Noise, Device and Room Robustness Methods for Pronunciation Error Detection
Ville-Veikko Eklund, Aleksandr Diment, Tuomas I. Virtanen · 2022 30th European Signal Processing Conference (EUSIPCO) · 2022
In this work, we address the problem of audio classification operating on signals recorded with various mobile devices in challenging environments. We propose a method for device, room and noise robust pronunciation error detection. It involves a data augmentation pipeline of convolution operations with room impulse responses and mobile device microphone impulse responses, and addition of background noise. A dataset of impulse responses of a diverse set of mobile devices, rooms and noises is collected. The method is evaluated in a pronunciation error detection task. The data consists of Finnish people uttering various English words accompanied by expert annotations of pronunciation errors. Classification accuracy is shown to improve by up to 12.9 percentage points as the amount of generated training data is increased. Given the large diverse set of collected impulse responses, we demonstrate that robustness is achieved consistently for new rooms and devices, excluded from the training set.