Video Deception Detection through the Fusion of Multimodal Feature Extraction and Neural Networks

Yuanya Zhuo, Vishnu Monn Baskaran, Lillian Yee Kiaw Wang, Raphaël C.‐W. Phan · 2024

Detecting deceptive behavior in videos is a complex task within several domains, including academic fraud assessment, commercial anti-fraud activities, judicial system evidence analysis, suspicious activity detection in security monitoring systems, and behavioral intent analysis in psychological research. In this study, we present a novel approach to video deception detection by integrating visual and audio models for deep feature fusion, primarily targeting advanced deception detection datasets. Our visual model leverages hierarchical image feature learning to enhance deceptive cue detection, complemented by an audio model that processes acoustic signals for precise speech pattern analysis. This multimodal method significantly boosts detection accuracy and lessens reliance on extensive training data. Notably, our visual model incorporates knowledge distillation technology, improving efficiency and reducing computational resource needs without compromising performance. We implement a transformer architecture using distillation tokens for effective learning and incorporate convolutional neural network insights to enrich our model’s interpretative capabilities. Experimental results demonstrate that our approach surpasses existing technologies in various standards and scenarios, offering enhanced deception recognition capabilities and addressing the challenge of limited training data.

Read the paper · More papers on PaperTik