Comparison of neural architectures for sensor fusion

Barbara Helga Talle, G. Krone, Gerald John Balm · 2002

For technical speech recognition systems as well as for humans it has been shown that the combination of acoustic and optic information can enhance speech recognition performance. But it still remains an open question, at which stage of processing the two information channels should be combined. We systematically investigate this problem by means of a neural speech recognition system applied to monosyllabic words. Different fusion architectures of multilayer perceptrons are compared both for noiseless and noisy acoustic data. Furthermore, different modularized neural architectures are examined for the acoustic channel alone. The results corroborate the idea of separate processing of the two channels until the final stage of classification.

Read the paper · More papers on PaperTik