V2Coder: A Non-Autoregressive Vocoder Based on Hierarchical Variational Autoencoders
Takato Fujimoto, Kei Hashimoto, Yoshihiko Nankaku, Keiichi Tokuda · IEEE Access · 2025
This paper introduces V2Coder, a non-autoregressive vocoder based on hierarchical variational autoencoders (VAEs). The hierarchical VAE with hierarchically extended prior and approximate posterior distributions is highly expressive for modeling stochastic components of speech waveforms. V2Coder learns the stochastic components as hierarchical latent representations in a data-driven manner and generates diverse waveforms in the time domain. VAEs tend to suffer from a phenomenon known as posterior collapse, in which little data information is encoded in the latent variable. To address this problem, we introduce a carefully designed architecture and skip loss that encourage encoding to latent variables in deep layers. Additionally, VAEs suffer from low-quality samples generated using prior distribution due to the prior hole problem. To improve the sample quality, we impose a constraint on latent information in each layer of the hierarchical VAE and demonstrate that this constrained optimization significantly affects the sample quality. Experimental results using single-speaker and multi-speaker corpora showed that V2Coder is a competitive neural vocoder comparable to existing non-autoregressive neural vocoders based on deep generative models. V2Coder generates high-quality speech waveforms faster than real-time on both GPU and CPU.