Wav2f0: Exploring the Potential of Wav2vec 2.0 for Speech Fundamental Frequency Extraction
Rui Feng, Yin-Long Liu, Zhen-Hua Ling, Jiahong Yuan · 2024
Speech fundamental frequency (F0) extraction is one of the most important tasks in speech signal processing. This paper aims to explore the feasibility of using deep learning for speech fundamental frequency extraction. Our approach, Wav2f0, combines the Wav2vec 2.0 model with fully connected and LSTM layers, leveraging the pretrained representations learned by Wav2vec 2.0. We conduct training and evaluation on a Vietnamese tone production corpus, which contains parallel recordings of Electroglottograph (EGG) and microphone signals. Wav2f0 outperforms Praat in pitch extraction accuracy on the corpus, especially in scenarios where Praat fails to estimate pitch. The Gross Pitch Error (GPE) of Wav2f0 is 7.1%, representing a more than 50% error reduction compared to Praat's 15.5%.