Background Music Removal Using Deep Learning
Shun Ozawa, Tomomi Ogawa · 2023
Due to the increasing popularity of video-sharing websites and social networking services, videos are frequently posted to the Internet. However, users must be careful when posting copyright materials, e.g., music. Thus, the purpose of this paper is to investigate a method to remove music from audio data that include a mixture of various sounds, e.g., driving noises, human conversation, and music. Although speech enhancement techniques can be utilized to remove music, such methods also remove noise other than music, which reduces the realism of the content. Sound source separation can also be utilized for music removal, however, it has to select non-music sounds after separation, which increases the process. Thus, we conducted a preliminary experiment using a deep learning method in order to remove only music from audio data that include a mixture of speech and music sounds in a monaural audio signal, maintaining original noises and without sound source separation. We used Conv-TasNet for the network structure and the parameters were updated such that the RMSE value between the output of the model and the original speech sound was small. The results showed that accuracy tended to improve with longer sample lengths in the range of 50–1,000 ms. The music removal results demonstrated that music components were reduced overall.