Experimental study of impression-based statistical mapping between speakers' faces and their voices
Yasuhito Ohsugi, Daisuke Saito, Nobuaki Minematsu · The Journal of the Acoustical Society of America · 2016
Recently, many types of voice-user interface (VUI) have been developed and some of them become prevalent. To realize more natural conversation between a user and a machine, one feasible approach is expected to be enhanced personification of the machine. In this study, we attempt to give an adequate face to the voice quality of the speech synthesizer embedded in the machine. A research question here is what kind of face should be provided to a given voice quality? Also, an inverse question is possible when a face is given in advance: what kind of voice quality should be mapped to a given face? To solve a problem of mapping between a set of elements and another set, a statistical mapping is often applied, where a parallel and pairwise corpus of elements and their corresponding ones is required. In this study, GMM-based mapping is adopted as statistical mapping between voice features and face features. Further, a parallel corpus is prepared by asking subjects to select suited voices (faces) for given faces (voices) based on their impressions on the faces and the voices. Experiments showed that, although adequate voices can be automatically generated for given faces, mapping between them was strongly subject-dependent.