ConvNeXt Based Deep Neural Network Model for Detection of Deepfake Speech

Özkan Arslan · 2024

Deepfake speech, generated by artificial intelligence or through the manipulation of existing speech recordings, raises serious concerns in areas such as privacy, social security, identity verification, and media integrity, posing an increasingly significant threat. Therefore, the development of speech-based deepfake detection and classification models is of paramount importance. This study employs a hybrid approach that integrates the ConvNeXt-Tiny model with machine learning-based techniques to identify and detect deepfake speech. The pre-trained ConvNeXt-Tiny model, which demonstrates effective performance in balancing detection accuracy and speed, is applied to speech samples obtained from the Fake-or-Real dataset, and deep features are extracted. The extracted deep features are then transferred to Extreme Learning Machine (ELM), Support Vector Machine (SVM), Random Forest (RF), and Deep Neural Network (DNN) models. Experimental results show that the hybrid structure of ConvNeXt-Tiny deep features based on transfer learning and DNN achieves an accuracy of 96.64%, an F1-score of 96.63%, and a Cohen’s Kappa value of 0.893, outperforming other models. The proposed framework presents a model with accuracy comparable to state-of-the-art methods for detecting scenarios where an adversary transmits deepfake speech through channels such as phone calls, voice messages, or similar audio communication methods.

Read the paper · More papers on PaperTik