McRTSE: Multi-channel Reverberant Target Sound Extraction
Shrishail Baligar, Shawn D. Newsam · 2024
We address the challenge of target sound extraction and detection in reverberant environments for general sound classes. We highlight the limitations of existing approaches when applied to reverberant conditions. In response, we develop an end-to-end Multi-channel Reverberant Target Sound Extraction (McRTSE) model, that outperforms the current methods in these conditions. Recognizing the need for realistic training datasets, we introduce the rNIGENS and rNIGENS-XL datasets, which accurately reflect real-world reverberant mixtures. We then use dual- and triple-channel versions of these datasets to train and evaluate our solutions. McRTSE is robust in handling challenging reverberant audio soundscapes and outperforms the baseline in separation and detection performance.