Multi-Resolution Singing Voice Separation

Yih-Liang Shen, Ya-Ching Lai, Tai-Shih Chi · 2024

It has been shown that time-domain neural networks achieve higher performance than networks working on the short-time Fourier transform (STFT) domain. The fully-convolutional time-domain audio separation network (Conv-TasNet) is an end-to-end separation model with outstanding performance. However, the fixed convolution kernel length in Conv-TasNet implies that it analyzes signals using the frequency resolution constrained by the kernel length. This paper proposes a multi-frequency-resolution (MF) architecture, which analyzes sig-nals using more frequency resolutions, and compares the MF model with Conv-TasNet on singing voice separation. The results show that the MF architecture improves performance of Conv-TasNet. In addition, we also demonstrate the MF architecture does not provide consistent benefits to the STFT-domain separation model.

Read the paper · More papers on PaperTik