Audiovisual speech enhancement: new advances using multi-layer perceptrons

Laurent Girin, Laurent Varin, G. Feng, Jean‐Luc Schwartz · 2002

This paper deals with the improvement of a noisy speech enhancement system based on the fusion of auditory and visual information. The system was presented in previous papers and implemented with a simple stimuli corrupted with white noise. Its principle consists of an analysis-enhancement-synthesis process based on a linear prediction (LP) model of the signal: the LP filter is enhanced thanks to associative tools that estimate the LP cleaned parameters from both noisy audio and lip shape information. The structure of the system is reviewed and we focus on the improvement that concerns the associators: multi-layers perceptrons are used instead of linear regression. It is shown that in the context of VCV transitions corrupted with white noise, the performances of the system are improved in terms of the intelligibility gain, distance measures and classification tests.

Read the paper · More papers on PaperTik