Research on perceptual fusion of audio and video based on deep learning
Qing An, Yanhua Chen, Shusen Wu · 2020
In view of the technical problems existing in the perception and fusion of unstructured data such as audio and video, the method of deep learning is used to realize the pixel level sound source location of video by combining the synchronization of audio and video content in the physical scene and analyzing the sound and image jointly. At the same time, aiming at the problem of low resolution and poor quality of face image in video, a recognition method based on face super-resolution is proposed to realize the identity attribute calculation of the target person.