A T-F Masking based Monaural Speech Enhancement using U-Net Architecture

Khadija Akter, Nursadul Mamun, Md. Azad Hossain · 2023

In a real-world environment, the intelligibility and quality of speech are reduced inevitably when it is encountered noises. Speech enhancement aims to have reconstructed clean speech by suppressing unwanted ambient noise. Numerous types of research have been accomplished for this enhancement t3ask; some of them uses spectral mapping technique but fails somewhere in real-life condition. This study proposes a time-frequency (T-F) masking-based speech enhancement approach which resembles the human auditory peripheral subsystem using the ratio of clean and noisy signal magnitudes using a U-Net model. The proposed work is carried out in several seen and unseen noisy conditions with several SNR values. To assess the performance of the proposed enhancement approach, speech intelligibility, and quality scores using four objective scores have been evaluated. The proposed network showed improvement in terms of objective scores and spectral mapping-based method over state-of-art networks.

Read the paper · More papers on PaperTik