Enhancement of Esophageal Speech Using Statistical Voice Conversion
Hironori Doi, Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, Kiyohiro Shikano · Institutional Repositories DataBase (IRDB) · 2009
This paper presents a novel method of enhancing esophageal speech based on statistical voice conversion. Esophageal speech is one of the speaking methods for total laryngectomees. Although it allows laryngectomees to speak by generating a sound source and articulating it to produce audible speech sounds using their esophagus and vocal organs, the generated voices sound unnatural. To improve the naturalness of esophageal speech, we propose a voice conversion method from esophageal speech into normal speech (ES-to-Speech). A spectral parameter and excitation parameters, such as F0 and aperiodic components, of normal speech are separately estimated from the spectral parameter of the esophageal speech in the sense of maximum likelihood using different Gaussian mixture models. We conduct objective and subjective evaluations of the proposed method. The experimental results demonstrate that the proposed method yields significant improvements in naturalness of esophageal speech while maintaining its intelligibility.