Automatic speech synthesis unit generation with MLP based postprocessor against auto-segmented phoneme errors

Eunyoung Park, Sanghun Kim, Jae Ho Chung · 2003

The work presented is about the postprocessor, which improves the performance of an automatic speech segmentation system by correcting the phoneme boundary errors. The proposed postprocessor reduces the range of errors in the auto labeled results that are ready to be used directly as synthesis unit. Starting from a baseline automatic segmentation system, our proposed postprocessor trains the features of hand labeled results using a multi-layer perceptron (MLP) algorithm. Then, the auto labeled result combined with the MLP postprocessor determines a new phoneme boundary. For phonetically rich sentences, we have achieved 19.9% improvement for the frame accuracy, comparing with the performance of a conventional automatic labeling system. Also, we have reduced the absolute error about 28.6%.

Read the paper · More papers on PaperTik