Vae-Space: Deep Generative Model of Voice Fundamental Frequency Contours
Kou Tanaka, Hirokazu Kameoka, Kazuho Morikawa · 2018
Modeling the speech generation process can provide flexible and interpretable ways to generate intended synthetic speech. In this paper, we present a deep generative model of fundamental frequency (F0) contours of normal speech and singing voices. The generative model we propose in this paper 1) is able to accurately decompose an F0contour into the sum of phrase and accent components of the Fujisaki model, a mathematical model describing the control mechanism of vocal fold vibration, without an iterative algorithm, and 2) can represent/generate F0contours of both normal speech and singing voices reasonably well.