VIDEO SPEECH SYNTHESIS A FLEXIBLE INTEGRATED SYSTEM FOR ANALYSING AND SYNTHESISING SPEECH PRODUCTION IN THE VISUAL DOMAIN

N. Michael Brooke · 2024

Earlier papers (l,Z) have suggested the potential of a reproducible, controllable and well-specified visual stimulus for analytical investigations of speech perception by normal hearing subjects in the bimodal, audio-visual domain.The output of such a 'video speech synthesiser' would be an animated real time graphical display of the essential facial topography during Speech utterance production.This would be synchronised with the output of an audio speech synthesiser to generate the audio-visual stimulus.The kinds of experiments which may be undertaken depend upon the sophistication of the video synthesiser.For example, discrimination experiments to explore the nature of categorical per-:eption in certain bimodal phonetic spaces such as the /ma,ba,pa/ continuum,(3) may be carried out using a simple terminal-analogue version of the video speth synthesiser.More complex experiments using purely visual and coherent or conflicting audio-visual stimuli (h) require a more detailed control of the vishal articulatory movements and timings.A sufficiently well-specified video synthesiser might ultimately be used to help rehabilitate the hearing-impaired by providing a device in which small but significant variations could be con trolled and exaggerated during a period of perceptual training.The use of a computer offers the most flexible approach to the implementation of a video speech synthesiser.Implementation of a simple terminal-analogue model of facial articulation has been described previously (I), in which the

Read the paper · More papers on PaperTik