Reconstruction of Temporal Envelope using 1D FLCNN-LSTM Speech Enhancement model

M. Anitha, B. Anitha Vijayalakshmi, N.S. Bhuvaneswari · 2023

The primary objective of speech enhancement algorithm is to enhance speech perception characteristics that have been damaged by background noise. Ensuring speech quality and intelligibility is essential in speech processing analysis for effective communication. To do this, a new temporal modulation processing speech enhancement approach is developed that reconstructs the temporal envelopes (TEVs) in the time frequency (T-F) domain using a one-dimensional feature learning convolutional neural network and long short-term memory (ID FLCNN-LSTM). This deep learning architecture is used to restore distorted envelopes. This approach is tested by utilising 2,700 words from nine different speech samples combined with babble noise and SSN (speech spectrum shaped random noise) at various SNRs. The proposed algorithm's performance is assessed using the Short time objective intelligibility (STOI) and Perceptual evaluation of speech quality (PESQ) metrics. The suggested model improved STOI and PESQ scores as 0. S2 and 2. S0 respectively. The experimental outcomes predicted for quality and intelligibility scores outperforms TEV algorithm, 1D-CNN, and other existing models in speech enhancement under noisy conditions.

Read the paper · More papers on PaperTik