Estimating the acoustic characteristics of speech from visual speech signals

Ben P. Yuhas · The Journal of the Acoustical Society of America · 1988

Speech articulation produces acoustic and visual signals. When the acoustic signal is degraded by noise, the visual signal can provide compensatory information [W. H. Sumby and I. Pollack, J. Acoust. Soc. Am. 26, 212–215 (1954)]. The aim of the work to be presented is to explore the extent to which the transfer function of the vocal tract filter can be estimated from the visual speech signal. The approach was to use a neural network to obtain the short-term power spectral envelope of the acoustic signal given the corresponding visual image as input. The visual images were taken from laser disc recordings of speaker's faces [Bernstein et al., J. Acoust. Soc. Am. Suppl. 1 82, S22 (1987)]. A small box centered at the mouth was extracted and spatially subsampled. The output of the neural net was the acoustic spectral envelope of the corresponding acoustic signal obtained via the cepstral method. After being trained on 95 tokens of vowels and dipthongs, the network was tested on 24 images it had not seen. The net was able to estimate the spectral envelope associated with these images more accurately than more traditional approaches that used stored templates. [Work supported by AFOSR Contract No. 86-0246.]

Read the paper · More papers on PaperTik