Sequence-wise Optimization for Quasi-Harmonic Speech Waveform Modeling
Shaowen Chen, Tomoki Toda · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022
Quasi-harmonic models (QHMs) are effective meth-ods for representing a speech waveform with frame-wise param-eters and flexibly resynthesizing a speech waveform from them. The original QHM methods analytically extract those parameters by directly minimizing an error between resynthesized and original speech waveform segments frame by frame. However, such a frame-wise parameter extraction process suffers from information loss between individual frames, causing the quality degradation of a resynthesized speech waveform. In this paper, we propose a sequence-wise parameter optimization method based on back propagation (BP) by directly minimizing the reconstruction error of a whole speech waveform. The proposed method is capable of specifically compensating for the missing information between frames by making the parameter extraction process and resynthesis process consistent. We investigate the effectiveness of the proposed method by conducting experimental evaluations using real speech utterances. The experimental results demon-strate that the proposed method achieves a great improvement of the speech resynthesis quality, i.e., from 11.7 dB to 36.2 dB of the signal-to-reconstruction error ratio and from 0.83 to 0.99 of short-time objective intelligibility.