A semi-automatic method for transcription error correction for Indian language TTS systems

Swetha Tanamala, Jeena J. Prakash, Hema A. Murthy · 2017

Database for building text-to-speech (TTS) synthesis systems consist of text sentences and the corresponding spoken waveform. Errors in transcription is an important factor that leads to poor synthesis quality. Mismatches between recorded speech and the corresponding text transcriptions occur mainly due to addition, deletion and substitution of words or phrases in the text or speech data. Presently, such mismatches are identified and corrected manually after listening to individual speech waveforms. In this work, this process of transcription error correction is made semi-automatic by flagging off the error units at phone level, based on the log-likelihood values of each phone after forced Viterbi alignment. A user interface is developed that highlights the errors. The errors only need to be manually corrected. TTS systems are built for two languages with the complete dataset and with the dataset after removing the data with wrong units. Result of listening test conducted showed that there is a significant improvement in synthesis quality.

Read the paper · More papers on PaperTik