Noise Robust Cross-Speaker Emotion Transfer in TTS Through Knowledge Distillation and Orthogonal Constraint

Rui Liu, Kailin Liang, De Hu, Tao Li, Dongchao Yang, Haizhou Li · IEEE Transactions on Audio Speech and Language Processing · 2025

By cross-speaker emotion transfer (CSEF) in text-to-speech (TTS) synthesis, we synthesize speech for a target speaker with the emotion transferred from reference speech by another (source) speaker. Traditionally CSEF is achieved by decoupling speaker and emotion components in speech, that highly depends on the availability of clean reference speech. To address the above issue, we propose a novelNoise-robustCross-SpeakerEmotion Transfer TTS model, i.e. NCE-TTS. NCE-TTS integrates the noise-robust emotion information extraction and noise-robust speaker-emotion disentanglement into a unified framework with two new modules, namelyknowledge distillationandorthogonal constraint. The knowledge distillation aims to directly learn the emotion features of clean speech, from noisy speech, with a conditional diffusion model. The orthogonal constraint seeks to disentangle the deep emotion embedding and speaker embedding and further enhance the emotion-discriminative ability. Unlike the traditional cascaded approach of first denoising and then extracting features, we have built a new training framework that achieves better emotion transfer results in noisy scenarios. We conducted extensive experiments on a multi-speaker English emotional speech datasetESD. The objective and subjective results demonstrate that the proposed NCE-TTS can synthesize emotionally rich speech while preserving the target speaker's voice in various noisy scenarios, with a significant improvement compared to all advanced baselines.

Read the paper · More papers on PaperTik