CAT-Net: A Channel and Self-Attention TCN for Robust Frame-Level Overlapping Speech Detection

Yassin Terraf, Youssef Iraqi · 2025

Detecting overlapping speech is essential for improving the performance of speech processing systems such as speaker identification, diarization, and automatic speech recognition. However, existing methods often fail in noisy and reverberant environments, limiting their real-world applicability. To address this challenge, we propose CAT-Net, a novel lightweight architecture for robust frame-level Overlapping Speech Detection (OSD). CAT-Net combines multiscale channel-wise attention, which dynamically weights frequency bins based on their importance using parallel temporal convolutions. The weighted features are fed into a Self-Attention Temporal Convolutional Network (SA-TCN), which models short-and long-term temporal dependencies via dilated convolutions. To overcome the limitation of uniform weighting across the receptive field, a self-attention mechanism is applied after each dilated layer to adaptively reweight temporal features based on contextual importance. A final classification module labels each frame as overlapping or single-speaker speech. As part of this work, we construct two comprehensive single-channel OSD datasets: one derived from the GRID corpus for neutral speech, and another from the RAVDESS corpus for emotional speech. For each dataset, we generate multiple versions simulating clean, noisy, reverberant, and combined noise-reverberation conditions across a wide range of Signal-to-Noise Ratio (SNR) levels and noise types. This enables rigorous evaluation of OSD models under realistic acoustic environments. Experimental results demonstrate that CAT-Net outperforms state-of-the-art methods across all conditions while using significantly fewer parameters, underscoring its effectiveness, efficiency, and suitability for deployment in practical speech processing systems. Furthermore, integrating CAT-Net into a standard speaker diarization system results in consistent improvements across both clean and noisy conditions.

Read the paper · More papers on PaperTik