Audio-visual Fusion of Artificial Intelligence for Enhanced Human Recognition

Zening Li, Jiachuang Wang, Xiawei Yue, Fangyu Zhao, Xiaoling Wei, Tiger H. Tao, Nan Qin · Journal of Physics Conference Series · 2024

Abstract In complex rescue environments, the information that can be perceived by a single vision is often limited, and it is easy to miss the best opportunity to help disaster victims. We designed an intelligent robotic system that fuses audio-visual information to solve the problems of poor lighting, heat source interference, and line-of-sight occlusion in complex rescue environments. By carrying visible optics cameras, infrared cameras, and microphone arrays, the system realizes multi-modal information sensing of vision, temperature, and hearing in complex environments. And we designed an intelligent recognition algorithm for multi-modal fusion. It searches for rescuers outside the camera’s field of view through sound source localization, and uses audio-visual fusion for human recognition, which ensures the normal operation of the rescue system in obscured or poorly lit environments. We conducted data acquisition and testing experiments in six different complex environments. The results show an accuracy of 95.17% using the audio-visual fusion algorithm. It has better performance compared to the single visual modality in all scenes.

Read the paper · More papers on PaperTik