Exploring Receptance Weighted Key Value Model for Single-Channel Speech Enhancement
Yuanle Li, Yi Zhou, Hongqing Liu · 2024
Speech enhancement has significantly benefited from advancements in deep learning, particularly in terms of intelligibility and perceptual quality. Traditional time-frequency (TF) domain methods rely on convolutional neural networks (CNNs) or recurrent neural networks (RNNs) to predict TF masks or speech spectra. In this study, we introduce a novel network architecture, the Receptance Weighted Key Value (RWKV) model integrated with the Deep Complex Convolution Recurrent Network (DC-CRN), aimed at effectively training complex targets. This model combines the strengths of RNNs and Transformers, redefining the attention mechanism for speech enhancement tasks to avoid the exponential growth in computational complexity typically associated with traditional Transformer models. Through a series of experiments on the VoiceBank+Demand dataset, the RWKV-DCCRN model has demonstrated superior performance, efficiency, and scalability in handling long-range dependencies compared to traditional DCCRN models. Our implementation showcases the potential to optimize computational efficiency and performance in real-time speech enhancement applications, making it particularly suitable for environments with varying noise conditions and for devices requiring low latency processing.