Application of Multi-Scale Convolutional Time-Frequency Attention Mechanism in Complex Acoustic Echo Cancellation
Weidong Qin, Zihao Wu, Qiuyu Zhang, Yibo Huang · 2025
Time-frequency (T-F) attention mechanisms have significantly improved acoustic echo cancellation (AEC). However, data downsampling, often employed to reduce computational complexity, results in the loss of time-frequency information and diminished time resolution, severely affecting model performance and accuracy. To address these challenges, this paper proposes a Multi-Scale Time-Frequency Attention Network (MTFAN) to enhance energy distribution in the time-frequency domain, thereby improving echo cancellation precision and efficiency. The MTFAN module leverages multi-scale convolutions to extract features across different time and frequency scales. Employing a dynamic fusion strategy integrates time-domain and frequency-domain attention maps with the input features, allowing the model to prioritize important frequency components and time dynamics of echo signals. This approach enhances the accuracy of feature extraction and optimizes the signal's energy distribution. Furthermore, by combining MTFAN with convolutional layers, a Convolutional Feature Fusion module (CFM) is constructed, further enhancing the model's ability to extract and process signal features. Experimental results demonstrate that the proposed AEC technique exhibits superior performance in double-talk tests conducted in both clean and noisy environments. Specifically, the proposed method improves the Perceptual Evaluation of Speech Quality (PESQ) score by 13.23% compared to the existing F-T-LSTM approach. Moreover, it surpasses the DTLN-AEC method with a 0.28 increase in the Mean Opinion Score (MOS), further substantiating its superiority in improving user experience.