Speech-Visual Emotion Recognition via Modal Decomposition Learning
Lei Bai, Rui Chang, Guanghui Chen, Yu Zhou · IEEE Signal Processing Letters · 2023
It is becoming a mainstream feature fusion approach for speech-visual emotion recognition (SVER) by directly using neural networks to fuse the extracted speech and visual features. However, the heterogeneity between speech and visual modalities usually results in a distribution gap and information redundancy between the extracted speech and visual features, thus affecting the performance of the SVER. To this end, this paper proposes a SVER method based on the modal decomposition learning. It leverages the shared, private and reconstructed modal learning with a specifically designed loss to decompose the extracted speech and visual features into the shared and private subspaces to obtain the shared and private features, which effectively reduces the distribution gap and information redundancy between the extracted speech and visual features. Experiments on the BAUM-1s, RAVDESS and eNTERFACE05 datasets also show that the proposed method achieves a better result.