EX-Vector: Emotional X-Vector Transfer Learning for Speaker Recognition With Emotion Domain Adaption
Zhuo Yang, Dongdong Li, Yu‐Dong Cai, Zhe Wang, Hai Yang · IEEE Transactions on Audio Speech and Language Processing · 2025
In emotional speaker recognition, the emotion-mismatch problem arises due to the inconsistency of the speaker's emotional state between the registration utterance and the test utterance. In this paper, we propose a new embedding-based method, Emotional X-Vector (EX-Vector), to address the above problems through a transfer learning strategy. EX-Vector introduces an emotion domain classifier with a gradient reverse layer to transfer emotion embeddings to emotion-independent speaker embeddings for emotion domain adaptation. Furthermore, the method incorporates an embedding cross-mapping module, which explores how individual units of speaker features influence each other to enrich the speaker information in the embedding, supporting both transfer learning and speaker recognition tasks. Attentive speaker identity compensation adopts attention mechanism and residual structure to retain the part of emotion information related to speaker identity and compensate it into speaker embedding during speaker recognition. The method extends and compensates the embedding content in speaker intrinsic feature extraction and speaker recognition respectively, preventing information loss during embedding transformation. The proposed EX-Vector is evaluated on emotional speaker recognition corpuses Mandarin Affective Speech Corpus (MASC), Interactive Emotional Dyadic Motion Capture (IEMOCAP), Crowd-Sourced Emotional Multi-modal Actors Dataset (CREMA-D) and podcasts recordings database (MSP-PODCAST). EX-Vector shows significant improvement in the evaluation metrics of EER and DCF compared with state-of-the-art methods in the task of emotional speaker recognition.