Optimizing Speech Naturalness in Speech Synthesis Using Improved DCGAN
Dan Wang, Qingyang Su · Procedia Computer Science · 2025
Existing speech synthesis models have certain deficiencies in speech naturalness, especially in terms of sound quality, fluency and speech details, which often fail to meet the requirements of high-quality synthesis. To solve this problem, this study explores the application and optimization of the traditional Deep Convolutional Generative Adversarial Network (DCGAN) in speech synthesis. Due to the temporal and frequency domain characteristics of speech signals, the traditional DCGAN architecture needs to be adjusted for speech data to generate high-quality speech waveforms. The study improves the architecture of the generator and discriminator, using a one-dimensional convolutional layer to replace the traditional two-dimensional convolutional layer to adapt to the time series characteristics of speech. By taking the Mel spectrum as input, the generator learns to generate the frequency characteristics of speech, while the discriminator extracts the time-frequency characteristics of speech to make true or false judgments. In addition, this study also proposes an optimization strategy for the loss function, combining adversarial loss, speech reconstruction loss and feature matching loss to improve the naturalness and sound quality of the generated speech. Experimental results show that the performance of the improved DCGAN (with the introduction of L1 loss) has improved, with the average MOS score rising to 3.89. The naturalness and fluency have both improved to 3.8 and 3.9 respectively, indicating that L1 loss helps improve the quality of speech generation, especially in terms of speech fluency and intelligibility.