Multi-Step Quality-Oriented Training for Cross-Dataset Offline Iterative Speech Enhancement
Shih-Chuan Chu, Chung‐Hsien Wu · IEEE Access · 2025
In recent years, speech enhancement research has increasingly leveraged deep neural network models to improve performance. Although these models often achieve strong results on specific datasets, they typically require substantial memory and computational resources, limiting their practicality in real-world applications. Consequently, evaluating model performance across diverse datasets has become essential. In this work, we propose a novel SE framework, termed CDiffSE-AC, which integrates a diffusion model with Actor–Critic techniques to address the limitations of existing methods. Our approach processes noisy speech inputs using a diffusion model alongside a pre-trained large-scale language model, which serves as the Actor by generating labels for the noisy input. These labels are subsequently used as conditioning information for the SE model. A Critic module evaluates the enhanced output by scoring its speech quality. During training, the enhanced outputs are recursively reintroduced as inputs to both the Actor and SE models, enabling iterative refinement and performance enhancement. We evaluate the proposed method using two different-sized subsets of the TIMIT and VCTK-DM datasets. Although CDiffSE-AC yields marginal score improvements within individual datasets, it demonstrates substantial gains in cross-dataset evaluations, significantly outperforming the baseline MOSE method. Specifically, CDiffSE-AC achieves PESQ improvements of 0.15 (7.2%) and 0.18 (7.5%) on the TIMIT and VCTK-DM test sets, respectively.