Real Time Speech Enhancement Using Triple Attention U-Net
Maisevli Harika, Papana Krishna Reddy, Chaitanya Jannu, Veeraswamy Parisae, Chinta Venkata Murali Krishna, G. L. Madhumati · 2024
Recently, the attention mechanism has been explored for improving speech quality, showing significant improvements. The attention modules have been developed to improve CNN's backbone network performance. However, these attentions often use fully connected (FC) and convolutional layers, which increase the model's parameters and computational costs. The proposed structure involves a layered Temporal Convolutional Network (TCN) at the core, with each convolutional layer in both the encoder and decoder of the U-Net architecture being succeeded by a triple attention block. An S-TCN is incorporated between the encoder and decoder to capture long-range dependencies within speech. A Triple attention block (TAB) is crafted to enhance the model performance, allowing it to concurrently concentrate on important areas in the channel, spatial, and time-frequency dimensions. The proposed speech enhancement system is assessed against baseline deep learning techniques and yielded better results in terms of PESQ and STOI.