Singing Voice Separation Using Interpretable Deep Learning
Christos Filippidis · Zenodo (CERN European Organization for Nuclear Research) · 2022
Audio Source Separation concerns the field of study, where the general aim is to isolate the sources from an auditory mixture. Deep learning models, which are frequently used for audio source separation, have contributed towards significant improvements in recent years. However, their black-box nature might lead to unintended effects, such as reinforcement of biases, because of the difficulty of understanding their inner workings. Thus there has recently been an increasing interest in the development of models that provide explanations for their decisions. Given that there is a lack of research in interpretability in the audio domain, in this thesis we carry out a series of experiments to leverage an existing interpretable model designed for sound classification, to singing voice separation. By using the mechanisms provided by the interpretable nature of the model as well as by analyzing the model’s predictions through additional interpretable methods, we facilitate the masking process for the singing voice separation task. The masks generated by the current intrinsic interpretable explanations of the model are not suitable for carrying out source separation tasks because of the low resolution of the computed saliency masks. However, after analyzing the model with other visual and auditory post-hoc explanations, we achieved sharper results in saliency maps, which have been used as our masks and have resulted in a significantly improved separation of the vocals. We compute the SDR and the SDR improvement (NSDR) for the reconstructed vocals estimations, in order to evaluate the separation and discuss the results. Even though the results of the evaluation are not in line with the performance of other state-of-the-art source separation algorithms, our method offers a novel approach in the field of interpretability in the source separation field and the audio domain.