Simultaneous Speech Denoising and Super-Resolution Using mGLFB-Based U-Net, Fine-Tuned via Perceptual Loss
Hwai-Tsu Hu, H. Tsai · Electronics · 2025
This paper presents an efficient U-Net architecture featuring a modified Global Local Former Block (mGLFB) for simultaneous speech denoising and resolution reconstruction. Optimized for computational efficiency in the discrete cosine transform domain, the proposed architecture reduces model size by 13.5% compared to a standard GLFB-based U-Net, while maintaining comparable performance across multiple quality metrics. In addition to the mGLFB redesign, we introduce a perceptual loss that better captures high-frequency magnitude spectra, yielding notable gains in high-resolution recovery, especially in unvoiced speech segments. However, the mGLFB-based U-Net still shows limitations in retrieving spectral details with substantial energy in 4–6 kHz frequencies.