Exploiting Inconsistencies in Object Representations for Deepfake Video Detection

Kishor Kumar Bhaumik, Simon S. Woo · 2023

AI advancements in recent years have made it more simpler to fake digital images and videos, and much harder to spot forgeries. When spread maliciously, these forgeries not only threaten the integrity of our political systems but also present serious moral and legal quandaries. Deepfake videos are mostly generated frame-by-frame, leaving visible object-level inconsistencies in both temporal and spatial dimensions. In this paper, we propose a novel deepfake video detection method that exploits this important clue. Specifically, we extract object representations using vision transformers from video frames and then model the object-level coherence in both intra-frame and inter-frame manner. We propose that our model can learn object-level visual inconsistencies by excluding irrelevant context and focusing on key traces responsible for compromised video authenticity.

Read the paper · More papers on PaperTik