Capsule Networks and LSTM Models for Robust Deepfake Detection in Audio and Video
M B Abisha, G. Jaspher W. Kathrine, S. Kushmitha · 2024
Due to the fast advancement of deepfakes, notable difficulties have been presented by deepfake innovation in confirming the believability of computerized content, especially on the audio and video front. This paper introduces a new hierarchical model that articulates Capsule Networks within LSTM for the effective and robust deepfake detection in both audio-video frameworks. It is argued that using Capsule Networks, spatial mindfulness capabilities enable the detection of subtle spatial objects in video outlines, whereas, LSTM models capture transient objects in audio arrangements and video frames over time. Thus, our approach implements two designs at once that reached progressed discovery accuracy, addressing spatial and temporary disorders, inherent in deepfake media. It is discovered from the outcomes that this model is very much effective on several datasets suggesting its better generalizing capability in different types of Deepfake media. In its current form, the design of this approach seems capable of being used for the real-time detection task. In this regard, this research provides practical insights into the development of media forensics offering an approach that can be more effective in the role of dissemination of fake news prevention, and fills the gap in the current literature, thus advancing the overall understanding of media forensics.