Fast Neural Vocoder With Fundamental Frequency Control Using Finite Impulse Response Filters

Yamato Ohtani, Takuma Okamoto, Tomoki Toda, Hisashi Kawai · IEEE Transactions on Audio Speech and Language Processing · 2025

In practical use, a neural waveform generative model, which we call a neural vocoder, must be able to produce high-quality synthetic speech and perform real-time inference on a single CPU with flexible control of the fundamental frequency ($F_{0}$). To achieve this functionality, this paper proposes a novel neural vocoder called FIRNet, which is based on the source-filter model. The basic concept of FIRNet is that multiple finite impulse responses (FIRs) are predicted from acoustic features using neural networks, and the speech waveform is then generated by filtering an excitation signal with these multiple FIRs. In the first version of FIRNet, a mixed excitation signal is employed, and the speech waveform is produced by convolving the mixed excitation with residual and resonance FIR coefficients. Although this FIRNet can achieve fast inference and$F_{0}$control, speech quality is not as high as that of other neural vocoders based on the source-filter model. To improve this, we apply three extensions to FIRNet: (1) a pitch-dependent convolutional network, (2) a structure separating the periodic and aperiodic components, and (3) data augmentation or multiple-speaker training. The experimental results show that the proposed method can improve speech quality while retaining flexible$F_{0}$control and fast inference.

Read the paper · More papers on PaperTik