Synthetic Speech Detection Using Deep Neural Networks

Irina Mutica, Şerban Mihalache, Dragoş Burileanu · 2024

In recent years, the proliferation of synthetic speech has raised concerns regarding its potential misuse for unethical activities including voice impersonation and deep fake generation. Addressing this challenge requires robust methods for detecting synthetic speech, which often exhibits subtle but discernible differences from natural speech. In this paper, three approaches for synthetic speech detection are proposed, two based on deep neural networks (DNNs), namely multilayer perceptrons (MLPs), convolutional neural networks (CNNs), and one based on an EfficientNetV2 model and transfer learning. The proposed system was trained on the Fake-or-Real (FoR) dataset, comprising utterances generated by some of the latest speech synthesis algorithms, and is able to generalize well on unseen samples generated with algorithms not encountered during training, yielding a validation accuracy of 98.9% and a test accuracy of 83.9%.

Read the paper · More papers on PaperTik