Sequence-Wise Speech Waveform Modeling via Backpropagation Optimization of Quasi-Harmonic Parameters

Shaowen Chen, Tomoki Toda · IEEE Transactions on Audio Speech and Language Processing · 2025

Sinusoidal modeling (SM) has been studied for speech modeling, which aims at representing and flexibly resynthesizing the speech waveform by the frequency and complex amplitude of several components. However, SM methods model the speech frame by frame, which inevitably leads to a loss of information between individual frames. In addition, SM method cannot adaptively optimize the prior frequency during the modeling. Even quasi-harmonic models (QHMs), which can adjust the frequency adaptively, require several iterations to update the frequency and finally obtain a dubious value. These limitations cause the quality degradation of a resynthesized speech waveform, especially in unvoiced speech parts. In this paper, we propose a sequence- wise speech waveform modeling method based on gradient descent to obtain the required parameters by directly minimizing the reconstruction waveform error and novel spectrogram loss, which can accelerate the convergence, of the entire speech between the result and the target. The proposed method can specifically compensate for errors caused by the missing information between frames by making the parameter extraction and resynthesis processes consistent. To investigate the effectiveness of the proposed method, we conduct experimental evaluations using real speech utterances, demonstrating that the proposed method achieves an improvement in speech resynthesis quality, i.e., from 10.6 dB to 14.9 dB of signal-to-reconstruction error and from 2.30 dB to 2.04 dB of mel-cepstral distortion.

Read the paper · More papers on PaperTik