Efficient Sequence Modeling in Time-Domain Speech Enhancement Using MP-RNN
Himani Daulat, Tarun Varma, Krishna Chauhan · 2025
Recent advancements in deep neural networks for speech enhancement emphasize the advantages of time-domain methods over traditional time-frequency techniques. However, processing long input sequences remains challenging due to computational limitations. The Multi Path-Recurrent Neural Network (MP-RNN) architecture overcomes this by partitioning long sequences into smaller, overlapping segments, enhancing computational efficiency and reducing optimization complexity. By integrating MP-RNN into the Time-domain-Audio-Separation-Network (Tas-Net), replacing conventional neural network modules, significant performance gains are achieved, with a reduction in model size by approximately 11 times on the WSJ0-2mix dataset. The performance is further evaluated under reverberant and noisy conditions, demonstrating the robustness of the MP-RNN architecture. Spectrogram and Power Spectral Density (PSD) analyses reveal noticeable improvements in signal clarity, with reduced noise and enhanced frequency representation in the processed signals. MP-RNN achieves a 7.89% improvement in Scale Invariant Signal-to-Noise Ratio (SI-SNR) and a 10.32% improvement in Signal-to-Distortion Ratio improvement (SDRi) over the Temporal Convolution Network (TCN). These results underscore MP-RNN's ability to efficiently handle long sequences, offering a scalable solution for real-world speech processing tasks.