Speech Enhancement for VHF Communication Audio Empowered by Feature Consistency and Contrastive Learning

Jinyi Gao, Weiwei Zhang, De Hu, Mingyang Pan · IEEE Transactions on Instrumentation and Measurement · 2025

Very-high-frequency (VHF) communication instruments are widely applied in vessels for transmitting both regular and emergency information. However, VHF communication speech often deteriorates severely due to complex transmission conditions. Most existing speech enhancement (SE) models ignore the correlation of different time-frequency (TF) units in the frequency domain or temporal samples in the time domain. Additionally, existing loss functions cannot effectively model spectral details with less intensities. As a result, speech partial loss and residual noise problems are serious for low signal-to-noise ratio (SNR) VHF communication speech. To address these problems, an SE network is proposed in this work. More specifically, a coarse-refine strategy is introduced to better disentangle speech and noise. In the coarse stage, the noise is preliminarily removed based on the encoder-conformer–decoder backbone and the consistency module. In the refine stage, the speech and noise partials are further disentangled based on consistency and contrastive learning modules. The consistency module makes the enhanced speech consistent with clean speech at both fine-grained and global levels, and the contrastive learning module can better distinguish clean speech from noise in hidden feature space. Experimental results demonstrate that the proposed network obtains a 28.0% higher perceptual evaluation of speech quality (PESQ) score and a 3.8% higher short-term objective intelligibility (STOI) score than the state-of-the-art (SOTA) network on the synthesized dataset. Also, it achieves a 4.5% higher mean opinion score (MOS) score than the SOTA network on the real-world dataset.

Read the paper · More papers on PaperTik