Applications of Deep Learning in Host Voice Synthesis and Emotional Expression
Yaxin Zheng · 2025
This paper proposes an end-to-end voice synthesis framework based on LSTM-GRU fusion model to solve the core problem of insufficient emotional expression in host voice synthesis. By integrating the long-term emotion modeling ability of LSTM and the local prosodic response advantage of GRU, combined with the hierarchical emotion modulation mechanism (32-dimensional emotion vector dynamically regulates the fundamental frequency and spectrum parameters), the collaborative optimization of the naturalness and emotional expression of the host's speech is realized. Experiments show that the model achieves 93.2% of emotional accuracy and 4.63 of naturalness (10% higher than optimal baseline) on the 200-hour professional host data set. The fundamental frequency trajectory error (5.3 Hz) and spectral distortion (4.1 dB) are significantly lower than those of the mainstream models (Tacotron2, FastSpeech2, etc.), especially in typical emotional scenes such as passionate and cordial hosting. The research provides a highly expressive and efficient solution for intelligent broadcasting system.