Speech Enhancement Based on Dual-Path Cross-Parallel Conformer Network

Qing Zhao, Ying Gao, Zhuoran Cai, Shifeng Ou · IEEE Access · 2024

The speech quality can be reduced by long-term and short-term noise, and frequency-domain single-path speech enhancement methods suffer from phase missing. To address these problems, this paper proposes a two-path cross-parallel Conformer network model for simultaneously modeling amplitude and complex frequency domain features. This model incorporates a residual connected cross-parallel Conformer module between the encoder and decoder. In this case, the self-attention mechanism Conformer module is incorporated into this model to extract local and global speech information in the time and frequency domains. Modules with cross-attention mechanism are then used for interaction in order to obtain more accurate information. The Attention Aware Fusion module is also added to this model. It aims to enhance the phase of the signal through complex spectrum reconstruction, solve the problem of phase information always being ignored, and fuse the amplitude and complex domain features learned from the dual path structure for optimal spectrum estimation. Afterwards, ablation experiments and other experiments are conducted on the Voice Bank + DEMAND database. The proposed speech enhancement model is evaluated using multiple metrics, including Perceptual Evaluation of Speech Quality (PESQ), Segmented Signal-to-Noise Ratio (SSNR), Short-term Objective Intelligibility (STOI), and three Mean Opinion Score (MOS) metrics. Moreover, the DPCPCNet model demonstrates outstanding performance in speech enhancement tasks, achieving a PESQ score of 3.00 on the dataset. The obtained evaluation indicators show that the proposed speech enhancement method outperforms other existing methods.

Read the paper · More papers on PaperTik