Deep Inception V5 Convolution Neural Network as Latent Features Supervision Network Towards Deep Fake Detection
Shanmugam Narmatha, S. Mythili · 2024
Deepfake has been widely exploited in recent years across various areas of social media platforms, movies, and news industries, which has led to multiple serious security concerns for society. Artificial intelligence approaches, in particular, make it easier to create and populate fake content on public platforms. Thus, it becomes mandatory to detect and prevent illegal deepfake content propagation. However, many researchers have developed Deepfake detection approaches using machine learning and deep learning architecture while those approaches bring more challenges in the form of generalization, underfitting and overfitting issues due to limited training labels. Hence, a new latent feature supervision network entitled Deep Inception V5 Convolution Neural Network increases the detection speed and detection accuracy with the general layer of the deep learning network on factorizing. Architecture provides excellent generalization on detecting the manipulated region of the content when processing the spatial and temporal features of the data. An inception network is highly capable of capturing a global and local distribution of the content effectively with the effective use of multiple filters simultaneously instead of increasing the layer of the network. Initially, content preprocessing is carried out to enhance the quality of the content based on transformation which changes the structure of the data and augmentation is to increase the structure of the data in terms of size and multiple characteristics. Those preprocessed data are projected to deep inception V5 architecture, which is composed of five inception modules with a convolution layer composed of filters of multiple sizes to capture the details of the image, The factorization layer is to reduce the parameter on factorizing the larger convolution to smaller convolutions, max pooling layer which calculate the maximum to the feature in the feature map and obtain the global features and local information separately, output layer provides the feature map and fully connected layer uses the SoftMax function through XG boost classifier to classifies the manipulated region from normal region accurately. And loss function through cross entropy to eliminate the overfitting and underfitting issues. Experimental analysis of the proposed deep inception V5 module is carried out using the DFDC dataset which is considered a deepfake dataset in a Python environment. Further performance of the proposed model is evaluated using cross-fold validation on test data against the conventional approaches. Finally, the proposed architecture provides 95.7% accuracy while training the model and 93.8% accuracy while validating the model.