Enhancing recognition systems through an integrated processing of visual and audio information
Eric O. Postma, Nikola Kirilov Kasabov, H.J. van den Herik · 2002
The AVIS framework for integrated audio and visual information processing is applied to the problem of person identification. An instantiation of the AVIS framework, called PIAVI, is based on a fuzzy neural network (FuNN) model of audio-visual person identification. In PIAVI's unimodal (visual) mode of operation, only dynamic visual features are used, whereas in the bimodal mode of operation, dynamic auditory and dynamic visual features are integrated at an early level of processing. Using a new dataset of dynamic features, a comparative study is performed with PIAVI in its two modes of operation. The results show that, with a large training set, perfect performance is achieved in the unimodal case. With a smaller training set, online application of the person identification system becomes feasible. Using this smaller set, unimodal identification performance is unsatisfactory. However, in the bimodal case, the identification performance is upgraded to satisfactory level of performance by early integration. It is concluded that, by using dynamic audio-visual features and FuNNs, an adequate on-line application of PIAVI in person identification tasks is within reach.