Single Channel Speech Enhancement using CNN with Frequency Dimension Adaptive Attention based Squeeze-TCN
Veeraswamy Parisae, S Nagakishore Bhavanam · 2024
In supervised speech enhancement, having context information is crucial for precise mask estimation or spectral mapping. The ability of generally used deep neural networks (DNNs) to capture temporal contexts is limited. Current research in speech enhancement predominantly concentrates on leveraging neural networks to efficiently capture long-term speech correlations to improve performance. Significantly, the distribution of energy in speech signals varies across different frequency channels. Therefore, for effective speech enhancement, it is essential for the model to prioritize and concentrate more on the informative frequency-specific feature channels. The proposed framework is an encoder-decoder structure comprising convolutional layers with S-TCN-FAA bottleneck to effectively capture long-range dependencies. In this approach, contexts are systematically aggregated through dilated convolutions present in squeeze temporal convolutional networks, which significantly expand the range of receptive fields. In contrast to LSTMs, temporal convolutional networks can achieve superior performance in modeling temporal sequences. Nonetheless, TCNs lacks the capability to represent the distribution of speech across the frequency dimension, which is of equal significance in maintaining the integrity of speech signals' structure. A module for adaptive attention in the frequency dimension within the context of the squeeze-TCN framework is introduced and used as a bottleneck layer in our proposed model. This module dynamically assigns varying weights to spectral components based on their importance in speech. The proposed model generalized well to untrained noises in our experiments.