A framework for parametric singing voice analysis/synthesis
Youngmoo E. Kim · 2004
The singing voice is the most variable and flexible of musical instruments. All voices are capable of producing the common phonemes necessary for language understanding and communication, yet each voice possesses distinctive qualities that are seemingly independent of phonemes and words. The unique acoustic qualities of an individual singer's voice arise from a combination of innate physical factors (e.g., vocal tract and vocal fold physiology) and time-varying characteristics of performance (e.g., pronunciation and musical expression). This research introduces a framework for singing voice analysis/synthesis that takes both physical and expressive factors into account by estimating source-filter voice model parameters (representing the physiology) and modeling the dynamic behavior of these features over time using a hidden Markov model (to represent aspects of expression). Historically, source and filter model features have been calculated independently, but here they are estimated jointly for better modelling of source-filter dependencies common in singing. Additionally, the vocal tract filter is estimated on a warped frequency scale, which more accurately reflects the frequency sensitivity of human perception. This framework has many possible applications, including singing voice analysis/synthesis and singer identification.