Xception Net & Vision Transformer: A comparative study for Deepfake Detection

Devanshu Shah, Dhiraj Shah, Dhruvi Jodhawat, Jinay Parekh, Kriti Srivastava · 2022

Deepfakes are artificial media in which an existing image consisting of a person is replaced with someone else. The creation of fake content is not new, but deepfakes are more credible because it involves the use of advanced deep learning techniques to manipulate or generate audio and visual content. In the 21st century, there have been astonishing advancements in Generative Adversarial Networks (GANs) and the use of encoder and decoder architecture [1] which have resulted in various effective methods of deepfake creation such as face-swap, lip-syncing, puppet-master and many more [2]. These methods are not only easily accessible but also get more accurate with time. Earlier deepfakes were detected with deep convolutional neural networks such as EfficientNet B7 and Xception Net, but with further advancements in deepfake generation there comes a need for better deepfake detection methods. This paper analyses the performance of Xception Net which has historically performed well with deepfake detection [3]–[5], Vision transformer [6] which is a fairly new technology and a combination of the two models. The paper proposes a combination of Xception and Vision transformer such that Xception Net is used for feature extraction from the patches and the output is then fed as a sequence to a transformer. This work is expected to assist readers to understand when to use Xception Net, Vision transformer and a combination of the two.

Read the paper · More papers on PaperTik