Multi-model Emotion Expression Recognition for Lip Synchronization
Salwa A. Al-agha, Hilal H. Saleh, Rana Fareed Ghani · 2020
Lip synchronization problem is a significant requirement in the multimedia world. The synchronization plan is in charge of guaranteeing that the sound and video streams are synchronized after handling. In recent years, there is a great deal of research, which was conducted on the issue of employing the recognition operation of emotion expressions, in the process of identifying and finding a solution to the problem of synchronization. This paper is based on the concept of recognizing the problem of synchronization in the multi-models audio/video of emotion expressions for non-synchronizing 3D video film stereoscopic. The proposed work is divided into two-stage: the recognition of audio expressions and the recognition of visual information expressions, with four basic emotion expressions: Anger, Happiness, Sadness, and Surprise. Multi-Layer Perceptron Back Propagation network is used for voice classification and depth image classification. Absolute Distance Differences is used for Intensity Frames of 3D Video (Geometric-based Feature) classification. The classification rate for the isolated audio signal is 90%. The classification rate for the geometric-based process of intensity image, plus depth image for different features of 3D video film is 80%. The classification rate for the geometric-based process of intensity image plus depth image for curvelet features of 3D video film is 82%. The recognition rate for the whole proposed system (audio plus video with curvelet features for depth image) is 80%.