Scat singing generation using a versatile speech manipulation system, STRAIGHT
Hideki Kawahara, Haruhiro Katayose · The Journal of the Acoustical Society of America · 2001
A set of procedures to generate scat singing by manipulating a small set of seed voices using a speech manipulation system called STRAIGHT [Kawahara et al., Speech Commun. 27, 187–207 (1999)] is proposed. F0 adaptive spectral smoothing based on a second-order cardinal spline combined with an F0 extractor based on a fixed-point analysis from filter center frequencies to the output instantaneous frequency enables the STRAIGHT system to generate a highly natural manipulated singing sound. Group delay manipulation for generating the excitation source signal introduces new control flexibility in source characteristics. F0 trajectories are generated using a dynamical model based on F0 feed-forward control regulated by auditorily mediated feedback [Kawahara et al., Vocal Fold Physiology, edited by H. Fletcher and P. Davis, pp. 263–278 (1996)]. Effects of the spectral interpolation function and interactions between F0 and the spectral envelope are also discussed and demonstrations of a scat chorus are presented. [Work supported by CREST and MEXT, Japan.]