Modeling audio-visual speech perception: back on fusion architectures and fusion control

Jean‐Luc Schwartz, Marie Cathiard · 2004

In a review paper about audio-visual (AV) fusion models in speech perception, we (Schwartz et al., 1998) proposed a taxonomy of models around two basic questions: architecture and control. Six years after, it appears that the proposals we made still seem rather convenient for discussing major questions about AV fusion. Moreover – and more importantly – recent experimental and theoretical progress seem to provide some elements of answer in both aspects. The aim of this paper is to review these elements, and to incorporate them into the general architecture-and-control framework. 1. FUSION ARCHITECTURES 1.1. The four architectures for audio-visual fusion In his well-known presentation of audio-visual models of speech perception, Summerfield (1987) introduced the concept of “metrics for audio-visual integration”, focusing on “the representations of the auditory and visual streams of information at their conflux”. The general literature on sensory interactions in cognitive psychology, and on sensor fusion in information processing, lead us conclude that there are four basic architectures (Schwartz et al., 1998). Their common point is that they should connect two separate inputs to one common output. The conception of a single output “loosing” in some sense the monosensorial nature of each input may be discussed (see the “convergence vs. association ” debate raised by Bernstein et al., in press). However, even in a conception of two separate routes in interaction from the input to the output, the questions addressed in this section remain valid, provided that they are rephrased in terms of: under what format are the A and V inputs represented in their sensory pathway when they interact in the route towards phonology or lexicon? In the Separate Identification (SI) model, the A and V representations are phonetic, that is mediated by the knowledge the subject has of his/her own language. In the Dominant Recoding (DR) model, the visual input is recoded into an equivalent sound, or into some of its spectro-temporal characteristics. In the Motor Recoding (MR) model, both the A and V inputs are in contact with a system analysing percepts in terms of the action able to have produced them. In the Direct Identification (DI) model, none of these process occur before phonetic identification which operates directly on a set of joined A and V parameters. In our view, the static vs. dynamic issue (or the shape vs. movement debate) is independent of the architecture. In consequence, the preference for static or dynamic parameters, if any, should not lead to select one or the other architecture. Notice that this position is itself controversial, and it has been often argued that movement and motor representation were linked topics (see e.g. Rosenblum & Saldana, 1998; Whalen et al., in press). However, we have several times advocated that recovering the vocal tract shape could be done without necessarily calling for dynamic features (Cathiard et al., 1996), and proposed a “shape from shading from movement” approach in line with recent neurophysiological computational models (Cathiard et al., 2003). 1.2. The “very early” route The four architectures share a common assumption of independence of the primitive monosensorial processing. That is, information would be first extracted separately in each sensorial channel before interaction and fusion. However, a number of recent studies on the detection of speech in noise have raised serious doubts about this assumption (since Grant and Seitz, 2000). Our own contribution was to determine if this gain in detection could contribute to a gain in identification. In a recent study (Schwartz et al., 2004), we showed, thanks to an original paradigm, that seeing the speaker’s lips does enable to better hear and hence better understand. The stimuli used in this set of experiments could not lead to lipreading per se since they corresponded to exactly the same lip gesture. However, intelligibility of these stimuli merged in noise was improved just because the acoustic cues were better extracted thanks to vision. The experimental trick consisted in dubbing the same lip gesture on a number of visually similar but auditorily different configurations, e.g. [y u ty tu ky ku dy du gy gu] in French. The visual stimulus did not enable to identify the syllable, but it provided a temporal cue improving the audio identification of these stimuli embedded in a large level of cocktail-party noise, and particularly the identification of plosive voicing. Replacing the visual speech cue (the lip rounding gesture) by a non-speech one with the same temporal pattern (a red bar on a black background, increasing and decreasing in synchrony with the lips) removed the benefit. Therefore, cross-modal interactions can occur early to enhance speech in noise and improve intelligibility. This indicates that there is, whatever the architecture, a preliminary set of interactions, that we called “very early” to make clear that they correspond to a contact point that should be distinguished from early interactions in the classical sense. This is likely to provide a number of interesting technological counterparts in terms of speech enhancement, source separation and audiovisual scene analysis (e.g. Girin et al., 2001; Sodoyer et al., 2002; Berthommier, 2003).

Read the paper · More papers on PaperTik