Concealing Audio Packet Loss Using Frequency- Consistent Generative Adversarial Networks

Linlin Ou, Yuanping Chen · 2022 5th International Conference on Pattern Recognition and Artificial Intelligence (PRAI) · 2022

Packet loss is one of the top reasons for speech quality degradation in Voice over IP calls. Packet loss concealment (PLC) is a technique of facing packet loss. This article describes a system based on a deep neural network (DNN) and pre-trained frequency-consistent generative adversarial network (Fre-GAN), which aims to heal audio impairments caused by packet losses, thus impairing the quality of the audio playout. We use the sliding window method to simulate real-time audio processing and extract features within the window, a multi-layer fully connected network is applied to predict the mel-spectrogram of missing audio packets based on the context within the window, and a frequency- consistent generative adversarial network is used to convert the mel-spectrogram to a waveform and backfill it to the part of the dropped frame. We name this system PLCfre-GAN. In the 2022 INTERSPEECH Audio Deep PLC Challenge, we applied PLCfre- GAN to mitigate artificially simulated audio impairments. The processed audio we submitted surpasses the challenge zero- padding baseline and ranks the top 5 with a PLC-MOS score on the blind test dataset of 3.478. Compared with previous machine learning methods, the proposed system has shown a considerable improvement in the scores of several evaluation metrics, and the recognition accuracy of synthesized speech is also guaranteed even in the case of frequent loss.

Read the paper · More papers on PaperTik