A Fast and Lightweight Speech Synthesis Model based on FastSpeech2

Huu-Kim Nguyen, Kihyuk Jeong, Hong-Goo Kang · 2021 36th International Technical Conference on Circuits/Systems, Computers and Communications (ITC-CSCC) · 2021

In this paper, we present a fast and lightweight speech synthesis model that is suitable for on-device applications. By leveraging the techniques of long-short range attention, depth-wise separable convolution, and linear attention, we significantly reduce the model size and complexity of the baseline FastSpeech2-based Transformer framework. Unlike the baseline model that requires O(N2) to compute attention and convolution operations because of nested-loop computations, our proposed model only requires O(N) computations due to the modification of a nested-loop into two cascaded single loops. Experimental results show that our proposed model is able to generate speech with a real-time factor of 0.26 and requires only 10.4 million parameters. Despite the reduction in model size and complexity, still, the generated speech quality of our model is nearly close to the baseline.

Read the paper · More papers on PaperTik