Attention-driven echo cancellation: A novel transformer-based approach for robust acoustic echo and noise cancellation
Soni Ishwarya V, Mohanaprasad Kothandaraman · Results in Engineering · 2025
Deep learning-based acoustic echo cancellation (AEC) systems have become increasingly popular due to technological advances and abundant data availability. They are well known for their outstanding performances and ability to deliver clear speech signals. Many deep learning models have been developed, including convolutional recurrent neural networks (CRNN), gated recurrent units (GRU), and bidirectional long short-term memory (BSTM). While these architectures have shown strong performance in acoustic echo and noise cancellation tasks, they may encounter limitations when handling longer sequences, such as vanishing or exploding gradients, and increased model complexity with higher computational requirements. This paper introduces a novel end-to-end deep learning framework called the fully transformer-based neural network (FTNN). The FTNN exclusively utilizes transformer architecture, incorporating only linear layers and attention mechanisms, omitting recurrent and convolutional components entirely. It incorporates two transformer-inspired attention mechanisms: single-head and multi-head attention. These attention mechanisms focus on intricate patterns, like those found in noisy doubletalk situations, where the nearend and farend speakers converse simultaneously amidst background noise. Its simplicity and efficiency characterize the proposed model. The model simplifies the architecture by relying solely on scaled dot product operations while enhancing efficiency. The FTNN's performance was evaluated using the TIMIT and AEC-Challenge databases, demonstrating its superiority over existing models in noisy doubletalk environments. It showed significant improvements in key metrics such as echo return loss enhancement (ERLE), perceptual evaluation of speech quality (PESQ), and correlation coefficient, proving its effectiveness in challenging acoustic scenarios. A case study was also included to evaluate the model's applicability. • Compact model suitable for real time implementation. • It relies entirely on transformer architecture, utilizing attention mechanisms. • It excludes LSTMs, GRUs, and CNNs, instead employing linear layers exclusively. • Its attention mechanisms enhance mask prediction accuracy. • It offers streamlined and effective training with only feed-forward layers.