A Dual-Branch Audio-Visual Event Localization Network Based on The Scatter Loss of Audio-Visual Similarity
Nan Shi, Xudong Yang, Lianhe Shao, Xihan Wang, Quanli Gao, Aolong Qin, Tongtong Luo, Qianqian Yan · 2023
In recent years, research on the audio-visual event localization has attracted much attention. Existing methods use audio-guided visual attention to lead the model to pay attention to the spatial area of the ongoing event, devoting to the correlation between audio and visual information but ignoring the correlation between audio, visual and spatio-temporal motion. In this paper, We propose a dual-branch audio-visual event localization network (DBAELN) that learns global and local event information on a sequence-to-sequence basis by combining audio and visual features as inputs at each period. And it uses a spatio-temporal feedback layer to provide fine-grained control for feature extraction, improving localization efficiency and accuracy. In addition, for fully supervised or weakly supervised settings, we propose a new scatter loss of audio-visual similarity (ASSDLOSS) to supervise feature learning, reduce the difference between generated samples and real samples, and finally obtain the best audio-visual event localization effect.