A study on acoustic features affecting speaker similarity between recorded voice and synthesized voice
Kenko Ota, Kohei Yoshida · The Journal of the Acoustical Society of America · 2016
This study focuses on the difference in the audibility between recorded voice and synthesized voice uttered by the same speaker. Voices whose pitch does not fluctuate in an utterance were recorded. Voices are synthesized by STRAIGHT. In this research, recorded voices and synthesized voices are allocated to a three-dimensional space by INDSCAL (Individual Differences Scaling) which is an analysis method of MDS (Multi-dimensional Scaling) in order to investigate the acoustic features which are related to the difference in the audibility between recorded voice and synthesized voice. Subjects compared following three combinations of voices: two recorded voices, two synthesized voices, a recorded voice and a synthesized voice. The acoustic features considered in this research are shown as follows: fundamental frequency, spectral centroid, spectral roll-off, spectral flux, spectral contrast, and spectral intensity. As a result, spectral roll-off and spectral contrast are related to the discrimination whether speakers of recorded voice and synthesized voice with same pitch are same or different. Moreover, fundamental frequency is also related to the discrimination whether speakers of recorded voice and synthesized voice with different pitch are same or different.