Pitch-invariant Speech Features Extraction for Voice Activity Detection

Ryhor Vashkevich, Elias Azarov · 2020

A new method for extracting the characteristic features of a speech signal that can be used in machine learning models to solve speech processing problems is proposed. The method is based on pitch-invariant convolutions on frequency axis of amplitude spectrum. On the example of voice activity detection, it is shown that the use of the proposed features makes it possible to get rid of excessive information content in the input data due to pitch invariance, which can significantly simplify neural network models, as well as reduce the amount of data needed for training. The proposed voice activity detection model is robust to a wide range of noises. The comparison with the publicly available voice activity detection model from the WebRTC showed higher F1 scores (0.94 vs 0.87).

Read the paper · More papers on PaperTik