DEEPFAKE ARTIFICIAL VOICE DETECTION. COMPARISON OF THE EFFECTIVENESS OF THE LSTM AND CNN MODELS.
А.Б. Абен, N.M. Zhunissov, G.N. Kazbekova, A.N. Amanov, A.A. Abibullayeva · NEWS OF THE NATIONAL ACADEMY OF SCIENCES OF THE REPUBLIC OF KAZAKHSTAN · 2025
This research presents a novel methodology for detecting DeepFake voices, which is based on the effective classification of fake and real audio signals. To enhance the assessment of information in the audience, the voices of 58 politicians and public figures were compiled as fake and real audio files. In the study, fake audio samples were artificially generated, while real samples were obtained from authentic sources. The analysis of the audio signal structure employed Mel-Frequency Cepstral Coefficients (MFCC), Zero-Crossing Rate (ZCR) metrics, and data visualization techniques, including bar charts and histograms. During the research, the numerical distribution, lengths, MFCC features, and ZCR values of the fake and real audio samples were analyzed. LSTM and CNN models were tested for DeepFake voice detection, resulting in the LSTM model achieving 100% accuracy, while the CNN model was rated at 97.50% accuracy. The findings demonstrated that the LSTM model could accurately and reliably distinguish between fake and real audio, emphasizing the importance of assessing the authenticity of audio signals in light of the dangers posed by DeepFake technology. This research provides functional methodologies aimed at developing systems for visual individuals while also uncovering new ways to determine the authenticity of audio signals and demonstrating the effectiveness of applying modern deep learning technologies. The study emphasizes that DeepFake plays a significant role in assessing and identifying information in an audience and provides a foundation for future research and practice.