RelUNet: Relative Channel Fusion U-Net for Multichannel Speech Enhancement

Ibrahim Aldarmaki, Thamar Solorio, Bhiksha Raj, Hanan Aldarmaki · 2026

Neural multi-channel speech enhancement models, specifically U-Net-based ones, show promising performance and generalization potential. These models encode input channels independently, and integrate them during later stages, or rely on extensive feature extraction methods. We propose a novel modification by incorporating relative information from the outset, where each channel is processed in conjunction with a reference channel through stacking. This exploits comparative differences to adaptively fuse information between channels, enhancing the representation of spatial features and improving the overall performance. Theoretical analysis shows that differential encoding leads to more compact manifolds, contributing to better generalization by capturing invariant features. The experiments conducted on the CHiME-3 dataset demonstrate improvements in speech enhancement metrics across various architectures. We show that our method enhances model capabilities in spatial feature processing, demonstrating strong performance in TDOA estimation on simulated datasets.

Read the paper · More papers on PaperTik