Voice Cloning Detection Using Deep Learning

Atef Isa Abdulla Husain Abdulla, Jenan Moosa · 2025

The rapid progress of artificial intelligence technologies has transformed how humans communicate with machines; voice recognition has been increasingly used as a unique “voice fingerprint” for authorization and task execution, motivating the development of technologies that use voice for identity verification. However, techniques have also emerged that attempt to falsify voices to impersonate others and invade privacy. Deepfake voice technology, which relies on creating data sets of human voices, uses deep learning and machine learning to analyze voice characteristics and build models that can replicate sounds with high accuracy. This research presents a deep learning-based model designed to detect forged or artificially synthesized voices. A dataset containing both genuine and synthetic human voices was used. Sound waves were transformed into frequency representations, then converted to digital matrices on which a Convolutional Neural Network (CNN) was trained. During training, the model extracted recurring features in sounds and classified them as real or fake. The main objective of this study is to evaluate deep learning's ability to achieve reliable results for security and authentication applications. The model achieved an accuracy of over 90% after training on the dataset. To further validate its effectiveness, the proposed model was compared to others, such as K-Nearest Neighbors (KNN), demonstrating its superior performance.

Read the paper · More papers on PaperTik