Improved weighted loss function for training end-of-speech detection models

Mikołaj Pudo, Adrian Wiśniewski, Artur Janicki · 2020

In this paper we propose an improved system for the detection of end of speech (EOS) events in noisy environments, needed, for example, in voice interfaces of mobile devices. Our solution is based on a deep neural network composed of convolutional, feed-forward and LSTM layers. For the input data we use mel-frequency cepstral coefficients (MFCC). The main novelty of our solution is the metric used during the training process of the model: our loss function returns higher values the later the model recognizes the EOS event. We confront this approach with the loss functions previously used, where such a delay was not considered. The experiments run on the TIMIT corpus, as well as additional evaluations on the other types of audio data, showed that our solution is significantly more robust to noisy and far-field environments compared to the baseline solution.

Read the paper · More papers on PaperTik