Analysis of RawNet2’s presence and effectiveness in audio authenticity verification

Karan Sreedhar, Utkarsh Varma, P. Balaji Srikaanth · 2024

Audio fabrication is on the rise around the world, mainly occurring in two different ways, spoofing and Deepfakes. Spoofing involves manipulating audio by editing and rerecording to make it appear however it is intended by the manipulator. Deepfakes are synthesized media in which a person is replaced with someone else's likeness, a rising phenomenon enabled by advances in deep learning and generative models. While initially focused on face swapping in images and videos, deepfake techniques have expanded to generating fake audio replicating a target speaker's voice. Deepfake audio presents risks such as facilitating fraud through mimicry of voices. This paper provides an analysis of RawNet2, an end-to-end model that has been used for both anti-spoofing and deepfake audio detection popularly, with the use of raw waveforms, and understand its relevance in modern audio authenticity verification, by tracking its effectiveness and evolution. ASVSpoof2019 and ASVSpoof2021 challenges were analyzed, and RawNet2’s usefulness across the years has been contextualized. As synthesis techniques continue improving, it is critical that detection methods also advance to counter the spread of disinformation. We outline promising directions such as leveraging self-supervision and meta-learning. This paper provides a timely snapshot of the deepfake audio arms race and tools to combat this emerging multimedia threat.

Read the paper · More papers on PaperTik