Knowledge enhancement for speech emotion recognition via multi-level acoustic feature
Huan Zhao, Nianxin Huang, Haijiao Chen · Connection Science · 2024
Speech emotion recognition (SER) has become an increasingly attractive machine learning task for domain applications.It aims to improve the discriminative capacity of speech emotion utilising a certain type of features (e.g.MFCC, Spectrograms, Wav2vec2) or multi-type combination features.However, the potential of acousticrelated deep features is frequently overlooked in existing approaches that rely solely on a single type of feature or employ a basic combination of multiple feature types.To address this challenge, a multi-level acoustic feature cross-fusion approach is proposed, aiming to compensate for missing information between various features.It helps to enhance the SER performance by integrating different types of knowledge through the cross-fusion mechanism.Moreover, multitask learning is utilised to share useful information through gender recognition, which can also obtain multiple common representations in a fine-grained space.Experimental results show that the fusion approach can capture the inner connections between multilevel acoustic features to refine the knowledge.The SOTA results were obtained under the same experimental conditions.