Deep Learning-Based Audio Splicing Forgery Detection Using PVTv2 Transformer
Elif Kanca, Tolgahan Gulsoy, Arda Üstübıoğlu, Beste Üstübioğlu, Selen Ayas, Elif Baykal Kablan, Güzin Ulutaş · 2024
Attacks on multimedia files by malicious users have become quite common, especially with the increase in the number of editing tools and their ease of use. Considering that such files can now be used both as evidence and for social visibility in all kinds of environments, it has become important to prove their authenticity. With the proposed method, the detection of merging forgeries in audio files has been carried out. For this purpose, the audio files received from the input are converted into cochleagram images. The PVTv2 based deep network architecture is trained with the generated cochleagram images. As a result of the training, the suspicious audio file given as input is labelled as original/fake. The proposed method gives 96.11% accuracy for 2s database, 94.63% accuracy for 3s database and 95.19% accuracy for 2s-3s database.