MesoNet-ViT: Meso Network and Vision Transformer for Deepfake Detection
İsmail İlhan, Mehmet Karaköse · 2025
Social media tools using deep learning technology are used for entertainment, but they can also pose a danger with forgery outside of their purpose. New solutions and detection methods are still needed in the challenge against forgery. In this study, the proposed MesoNet-ViT method was used to classify fake and real images. MesoNet-ViT is a combination of the Meso and Vision Transformer (ViT) methods used in deepfake detection. In this hybrid architecture, the face region is extracted from the input file with the BlazeFace method and given as input to the classification module. Feature maps extracted from the face image with the Meso method are fed to the ViT model, which detects whether the video is fake or real. When the experiments are analyzed, our method has shown that it can compete with other methods with an ACC value of 96.2 on the DeepFake Detection Challenge (DFDC) dataset and an ACC value of 98.1 on the DF-TIMIT dataset.