Prosodic and segmental factors in foreign-accent conversion
Daniel Felps, Heather Bortfeld, Ricardo Gutiérrez‐Osuna · 2008
We propose a signal processing method that transforms foreign-accented speech to resemble its native-accented counterpart. The problem is closely related to voice conversion, except that our method seeks to preserve the organic properties of the foreign speaker’s voice; i.e., only those features which cue foreign-accentedness are to be transformed. Our method operates at two levels: prosodic and segmental. Prosodic transformation is performed by means of time and pitch scaling. Segmental transformation is performed by convolving the foreign speaker’s excitation with the warped spectral envelope of the native speaker. Perceptual results indicate that our model is able to provide a 63 % reduction in foreign-accentedness. Multidimensional scaling also shows that the segmental transformation causes the perception of a new speaker to emerge, though the identity of this new speaker is three times closer to the foreign speaker than to the native speaker. Index Terms: voice conversion, foreign accent, speaker identity