Can Deepfakes Mimic Human Emotions? A Perspective on Synthesia Videos
Satota Mandal, Bratati Ghosh, Sunen Chakraborty, Ruchira Naskar · 2024
The rapid progress of deepfake technology has made it an increasingly significant topic in current times. While it offers valuable applications in various contexts, a major challenge lies in discerning deepfake contents. Deepfake refers to media content that is synthetically generated or altered using advanced AI techniques, many times with malicious intent to present it as genuine. This includes synthesis or manipulation of audio, video, images, and text to produce highly realistic but fabricated content. In this paper, we work towards identifying deepfake videos generated by advanced diffusion technology, by finding inconsistencies in emotional cues between real and synthetic human videos. In our experiments, we adopt Synthesia, an SOA diffusion model based synthetic video generator, vis-a-vis real videos collected from YouTube. To discern between genuine and synthetic videos, we follow a statistical approach which provides us strong cues in terms of emotional irregularities, and paves the path for detection of diffusion model generated deepfakes in the future. Alongside, we also develop a vision transformer model for fake video detection based on temporal variation in human facial expressions, which generated an accuracy of 98.11%. Our work proves that components of emotion constitute an essential trait for identifying deepfakes, strong and effective against even the most recent classes of deepfake generators.