Optimal Transport With Class Structure Exploration for Cross-Domain Speech Emotion Recognition
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Junhai Xu, Yongwei Li, Di Jin, Lin Zhang, Wenhuan Lu · IEEE Transactions on Audio Speech and Language Processing · 2024
Speech emotion recognition (SER) has widespread applications in human-computer interaction. However, the performance of SER models often drops in domain mismatch scenarios. Although existing unsupervised domain adaptation (UDA) algorithms could mitigate the domain mismatch problem by aligning feature distributions between domains, they do not consider the distribution relationship between emotion categories, thereby decreasing the capability to discriminate emotions. Moreover, directly aligning the probability distributions of different domains overlooks the emotional structure information contained in the target domain samples. To overcome this limitation, this paper proposes a novel UDA for cross-domain SER, termed Optimal Transport (OT) with Class Structure Exploration (OTCSE). OTCSE aims to measure the global probability distribution distance (GPDD) differences while considering the distribution relationship between emotional categories. Additionally, it explores the intrinsic structure information in the target samples. Specifically, we first embed joint OT into a deep SER framework to measure the GPDD between domains. Second, we propose to use a self-supervised learning (SSL)-based domain exploration module in the OT adaptation process to assist in exploring class structure information. Finally, we propose a two-step optimization strategy that allows OTCSE to update the model parameters of the SER and SSL modules end-to-end while solving the optimal transport coupling based on Sinkhorn's iteration algorithm. Experimental results in cross-domain SER demonstrate that the proposed OTCSE outperformed state-of-the-art UDAs. Further, we explore the complementarity of GPDD alignment and class structure exploration. This synergy alleviated the negative transport problem in OT and enhanced the efficiency of representation learning in SSL.