Artificial Intelligence to Combat Audio Fraud: A Flask-Deployed Hybrid Deep Learning System
Shashank Solanki, R. N. Ravikumar, Tushar Devataval, Jeet Gajera, Sushil Kumar Singh, Himanshu Gupta · 2024
As the use of synthetic communication is growing widely spread, the ability to accurately identify counterfeit spoken words is now more important than ever. In this study, we have developed an advanced system using deep learning techniques that can differentiate between real and modified audio files. Our research employs Kaggle’s DeepFake Voice Recognition Dataset and explores different deep learning architectures such as "Artificial Neural Networks (ANN)", "Convolutional Neural Networks (CNN)", and "Recurrent Neural Networks (RNN)". One crucial novelty of our work lies in merging these models with XGBoost, a machine learning algorithm, to create a hybrid system which greatly enhances the accuracy of detecting false audio tracks. What is surprising, however, is that RNN and XGBoost showed accuracy rate of 99.32% outperforming other settings. Additionally, we created an easy-to-use web application using Flask enabling users to upload and verify authenticity of voice records easily. This system tackles both the critical problem of voice forgery as well as advances in the field of sound authentication technology.