Audio-Visual Event and Sound Source Localization Based on Spatial-Channel Feature Fusion
Xiaolong Zheng, Ying Wei · 2022 7th International Conference on Signal and Image Processing (ICSIP) · 2022
In this paper we refer to scenes where visual and sound streams co-occur in video clips as audio-visual events (AVEs), and mainly focus on supervised and weakly supervised audio-visual event localization of unconstrained video clips and sound source localization in audio-visual events. Besides the commonly used spatial attention to learn the weight coefficients of the whole image, the channel attention to learn the weight coefficients of different feature channels is introduced. Key information of audio and video is fused in both spatial and channel domain. And then the fused features from the spatial attention part and channel attention part are further fused as the input of the Bi-LSTM for recognition. We conducted experiments on the public data set AVE, and compared with the previous methods, our method achieved new accuracy in both supervised and weakly supervised audio-visual event positioning, thus verifying the effectiveness of our method.