Audio signal Enhancement based on amplitude and phase deep learning model

Mahrez Benneila, Salah Hadji Line, Mouhamed Anouar Ben Messouad · 2024

In this work, we apply an approach based on simultaneous enhancement of the amplitude and phase time-frequency representation of a composite audio signal with two datasets, namely the Voice Bank+Demand and the NTCD TIMIT datasets. The approach is based on an encoder and a decoder connected by transformers having a convolution module. The encoder uses both amplitude and phase spectra as input, while the decoder has two parts one of parallel amplitude mask and the other is a phase decoder. Thus allowing direct recovery of clean amplitude spectra and also clean enveloped phase spectra, its activation function is a sigmoid learnable and it estimates the phase in parallel. Like all neural networks there are losses at several levels such as: amplitude, phase, and short-term complex, without forgetting the temporal waveforms. All these losses are used jointly in training the model implemented, the validation of this model is done on the two datasets and the experimental results gave a PESQ of 3.5 for the Voice Banktdemand and a PESQ of 3.22 for NTCD-TIMIT.

Read the paper · More papers on PaperTik