Audio Classification through MFCC features using RNN algorithm
K Bhargavi, Venna Deepa, I. Harshitha · 2024
Recurrent Neural Network (RNN) algorithm, is renowned for its effectiveness in object localization and classification tasks within the realm of computer vision. This study combines RNN capabilities with the analysis of Mel-Frequency Cepstral Coefficients(MFCC) features, which have demonstrated effectiveness in capturing unique audio signal characteristics. Mel-Frequency Cepstral Coefficients (MFCCs) serve as a compact representation of the spectral content of audio signals, enabling the identification of distinctive features crucial for detecting anomalies introduced during the generation of deepfake audio. The Recurrent Neural Network (RNN) algorithm is applied to these MFCC features. RNN excels in tasks such as image recognition, object detection, and semantic segmentation. Previous research in the field of audio forensics has explored various methodologies for detecting manipulated audio content. While some approaches focus on signal processing techniques, others employ machine learning algorithms. MFCCs have consistently demonstrated their efficacy in capturing audio features relevant to authenticity. The proposed approach enhances the detection of fake audio, a critical challenge in the era of advanced digital media manipulation. By combining the strengths of MFCC features and the RNN algorithm, the model aims to learn and identify intricate patterns associated with genuine audio, enabling accurate differentiation between authentic and manipulated audio sources. By adapting RNN to analyze MFCC features, the model gains the capability to discern intricate patterns associated with genuine audio, thereby distinguishing between authentic and manipulated audio sources.