Application of Machine Learning for the Detection of Audio Deep Fake
Divya K, Gunjan Chhabra, Pallavi, S.A. Tiwaskar, S. Hemelatha, Vivek Saraswat · 2024
Strong detection procedures are crucial for preserving digital safety and reliability in light of the rapid development of deepfaketechnology, especially in the audio domain. Presenting a "Improved Deepfake Audio Detection using Spectro-Temporal Deep Learning Approach," our study attempts to tackle the growing problem of differentiating real audio from advanced deepfake alterations. In order to train and test a new deep learning model, our approach makes use of the extensive ADD2022 dataset. At its heart, our method is based on a painstaking data preparation step that involves resampling, normalizing, and silence removal of audio samples to make them consistent and improve the quality of the model input. An novel combination of Convolutional Neural Networks (CNNs) for spectral feature extraction and Recurrent Neural Networks (RNNs) or Long Short-Term Memory (LSTM) networks for temporal dynamics analysis is used in our deep learning model architecture. The complex irregularities seen in deepfake audios may be easily detected by this hybrid model. This research presents a method for detecting speech spoofing using Convolutional neural networks. The system uses many audio properties to distinguish between human and synthetic voices. In the worst-case scenario, unethical applications of deepfake audios meant to harm third parties might do widespread damage to national assets and reputations. An assailant may create convincing voiceovers by using a tiny audio sample of a real person. A 2D graph may be created using mathematical formulas to represent any audio stream. Converting audio files to pictures of audio features (Spectrogram, MFCC, FFT, STFT) before logging obtaining the array values in a numerical format that is best for feeding into a convolutional neural network (CNN) reduces the amount of computation needed to build a system that can detect deepfake voices. Predictions are made using both standalone and combined methods of entering data into the model.