Advancing Limited Data Text-to-Speech Synthesis: Non-Autoregressive Transformer for High-Quality Parallel Synthesis
Mohammed Salah Al-Radhi, Omnia Ibrahim, Ali Raheem Mandeel, Tamás Gábor Csapó, Géza Németh · 2023
Despite the impressive results achieved by autoregressive generative models like Tacotron2 in end-to-end speech synthesis, their slow inference speed remains a significant drawback. To overcome this limitation, non-autoregressive Text-to-Speech (TTS) models like FastSpeech2 and neural vocoders like AutoVocoder, have emerged as faster alternatives with comparable quality. In this work, we present a novel lightweight Arabic TTS system based on a transformer architecture that utilizes fewer parameters than Tacotron2. Our approach combines convolutional and transformer-based blocks and is fully based on a non-autoregressive training framework. Our system can accurately reproduce the characteristics of natural speech like tone, pitch, timing and word pronunciation with state-of-the-art quality, making it suitable for practical applications like speech synthesis for low-resource languages and conversational agents. Our method is validated by acoustic analysis and subjective listening tests.