DEEPFAKE VOICE DETECTION USING ML

International Journal of Progressive Research in Engineering Management and Science · 2024

Recent advancements in AI-generated human voices have raised significant concerns regarding impersonation and the spread of disinformation, highlighting the need for effective methods to identify synthetic voices.This research presents a novel strategy for detecting synthetic human voices by identifying artifacts produced by vocoders in audio signals.Many DeepFake audio synthesis techniques utilize neural vocoders-neural networks designed to convert temporalfrequency representations, such as mel-spectrograms, into waveforms.By recognizing the processing effects of neural vocoders in audio samples, we can ascertain whether a voice is artificially created.To facilitate the detection of synthetic human voices, we propose a multi-task learning framework that employs a binary classification RawNet2 model, which integrates a vocoder identification module with a shared feature extractor.By framing vocoder identification as a pretext task, we guide the feature extractor to concentrate on the distinctive artifacts left by vocoders, thus enhancing the feature set available for the final binary classification.Our experimental results indicate that the enhanced RawNet2 model, which incorporates vocoder identification, demonstrates superior classification performance on the binary detection task.

Read the paper · More papers on PaperTik