AMH-Net: Adaptive Multi-Band Hybrid-Aware Network for Emotion Recognition in Speech
Hengrui Li, Yongbing Zhang, Shaohui Liu · IEEE Signal Processing Letters · 2025
Speech emotion recognition (SER) technology analyzes speech signals to automatically identify the speaker's emotional state. However, existing methods overlook feature extraction based on human acoustic characteristics. In this paper, we propose AMH-Net, an Adaptive Multi-band Hybridaware Network designed for SER. The model leverages formant (F1, F2, F3) from speech science, which describe the human vocal tract, to partition speech signals into multiple frequency bands. A variable-depth residual network structure is employed for more precise extraction of emotional characteristics. In addition, a hybrid attention mechanism is integrated to combine information, resulting in a more comprehensive emotional representation. Experimental evaluations of six diverse datasets show that AMH-Net outperforms state-of-the-art methods, achieving improvements of 2.11% and 2.64% in average UAR and WAR, respectively, on each corpus. The code is publicly available at https://github.com/hengruili1997/AMH-net