Model training using parallel data with mismatched pause positions in statistical esophageal speech enhancement
Mayumi Kishimoto, Tomoki Toda, Hironori Doi, Sakriani Sakti, Satoshi Nakamura · 2012
As one of the speaking aid techniques for laryngectomees, an esophageal speech enhancement method based on eigenvoice conversion has been proposed. In this method, conversion models are trained using utterance pairs of esophageal speech uttered by a laryngectomee and normal speech uttered by many normal speakers. Recording of normal speech of which pause positions correspond to those of esophageal speech is effective to develop well-designed training data for building the conversion models but it requires an enormous amount of time and expensive costs. In this paper, we propose a method capable of effectively using normal speech data including mismatched pause positions as training data by alleviating their impact on the conversion models. The experimental results demonstrate that the proposed method yields significant improvements in both speech quality and conversion accuracy for speaker individuality (i.e., speaker identity).