EnvFake: An Initial Environmental-Fake Audio Dataset for Scene-Consistency Detection
Hannan Cheng, Kangyue Li, Long Ye, Jingling Wang · 2024
In virtual reality scenarios, audio realism plays a crucial role in enhancing the user's immersive experience. Recently, many studies have focused on the determination of audio authenticity. However, existing studies mainly focus on the realism of the audio itself, e.g., automatic speaker verification, while ignoring the authenticity detection of the audio's background information. This background information usually contains ambient, reverberant, and reflective sounds, which can provide rich cues of scene-consistency detection. In this paper, we address the under-explored task of acoustic and visual scene-consistency detection by introducing a new Environmental Fake Dataset named EnvFake for the first step. EnvFake uniquely alters the acoustic scenes of recordings without changing the linguistic and visual content, providing a benchmark for training and evaluating models in this field. Our work shifts from traditional speech-centric forgery detection to encompassing a broader range of acoustic scenes.