End-to-end Visual-guided Audio Source Separation with Enhanced Losses
Duc-Huy Pham, Quang-Anh Do, Thanh Thi Duong, Thi‐Lan Le, Phi Le Nguyen · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022
Visual-guided Audio Source Separation (VASS) refers to separating individual sound sources from an audio mixture of multiple simultaneous sound sources by using additional visual features that guide the separation process. For the VASS task, visual features and the correlation of audio and visual play an important role, based on which we manage to estimate better audio masks to improve the separation performance. In this paper, we propose an approach to jointly train the components of a cross-modal retrieval framework with video data and enable the network to find more optimal features. Such end-to-end framework is trained with three loss functions: 1) separation loss to limit the separated magnitude spectrogram discrepancy, 2) object-consistency loss to enforce the consistency of the separated audio with the visual information, and 3) cross-modal loss to maximize the correlation of audio and its corresponding visual sounding object while also maximize the difference between the audio and visual information of different objects. The proposed VASS model was evaluated on the benchmark dataset MUSIC, which contains a large number of videos of people playing instruments in different combinations. Experiment results confirmed the advantages of our model over previous VASS models.