Sensory fusion: integrating visual and auditory information for recognizing speech
Gregory J. Wolff · 2002
A straightforward method of combining information from separate sources for pattern recognition is reported. Conditional class probabilities are computed for each channel independently, using a special network architecture. These probabilities are then combined according to Bayes' rule under the assumption of conditional independence. This method does not require any parameters other than those used to model each modality individually, is simple to compute, and automatically compensates for differences between the two channels. It is argued that this method can be a good starting point, even if conditional independence does not hold, especially when limited training data is available. Experiments in recognizing phonemes from acoustic and visual inputs indicate that this method can outperform more powerful models, and that is exhibits behavior similar to humans.>