Acoustic and linguistic effects in synthesized speech augmentation for speech recognition

Yohan Lim, Donghyun Kim, Sang Hun Kim · ETRI Journal · 2025

Abstract Recently, numerous studies have been conducted to incorporate the knowledge of massive text corpora into speech recognition via text to speech (TTS). However, the distribution mismatch between synthetic and real speech has always been an issue. In this paper, we analyzed how these mismatches affect the acoustic and linguistic aspects of automatic speech recognition (ASR) performance. For acoustics, we divided the acoustic mismatch into TTS‐related and non‐TTS‐related and analyzed how each acoustic mismatch affected ASR performance. Next, from a linguistic perspective, we experimented to determine how synthetic speech from a large text corpus affects the performance of speech recognition in various domains. The experimental results show that (i) substitution errors, which are the bulk of the recognition errors in ASR trained on synthetic speech data, are affected by the prosody mismatch between synthetic and real speech; (ii) pretraining ASR with synthetic speech data first and performing transfer learning with real speech outperformed training in the reverse order; and (iii) pretraining with a large amount of synthetic speech improves performance further in language model shallow fusion.

Read the paper · More papers on PaperTik