Unmasking Audio Deception: Performance Analysis in Machine Learning Models with Mel-Frequency and Gammatone Frequency Cepstral Coefficients

P P Sudharsana, R R Rajalaxmi, R Gughan, R. Thamilselvan, S. Saranya, Komirineni Sruthi · 2024

This study delves into the pervasive issue of audio deepfakes and their societal implications, emphasizing a Machine Learning (ML) approach for detection. Employing the Mel Frequency Cepstral Coefficients (MFCC) and Gammatone Frequency Cepstral Coefficients (GFCC), the research utilizes the FOR (FAKE or REAL) Dataset to evaluate the efficacy of various models, including XGBoost, Ada Boost, SVM, LightGBM, Logistic Regression, Extra Tree, Gaussian Naïve Bayes, Decision Tree, MLP Classifier, Linear Discriminant Classifier, and Quadratic Discriminant Classifier. The study critically assesses the potential of MFCC and GFCC features in discerning synthetic audio data and employs accuracy, precision, and recall as evaluation metrics. The results highlight nuanced distinctions in the performance of each model, shedding light on their respective capabilities in distinguishing genuine from deepfake audio content. This research contributes valuable insights into the ongoing efforts to fortify societal resilience against the escalating threat of audio deepfakes, offering a comprehensive evaluation of ML models and emphasizing the significance of MFCC and GFCC features in addressing this burgeoning concern [1]

Read the paper · More papers on PaperTik