Study on time-dependent voice quality variation in a large-scale single speaker speech corpus used for speech synthesis
Hisashi Kawai, Minoru Tsuzaki · 2004
The paper studies voice quality variation in a large-scale single speaker corpus used in recent corpus-based speech synthesis. First, a perceptual experiment is conducted to obtain scores for voice quality difference in a stimulus made by concatenating phrases collected from separate recording sessions. Second, acoustic measures are examined on their performance in classifying high and low scoring stimuli. Results show that band-limited power in the 8-16 kHz range performs best, closely followed by MFCC distance in the 0-4 kHz range, and that spectral tilts are almost irrelevant. However, the performance is not satisfactory for practical use (the equal error rate is 25%).