A Supervised Transformer-Based Model for Attributing Mobile-App-Generated Synthetic Audio Artifacts

Lucky Onyekwelu-Udoka, Yong Guan · 2024

The rising production and uses of synthetic audio in the online digital world, and various application domains have enabled many exciting applications. However, we also see the authenticity of audio artifacts (forensics) and their trustworthiness becoming severe problems in today's rapidly evolving digital world. We (forensic communities) have found that more and more such synthetic audio artifacts are generated using mobile apps. This paper focuses on detecting synthetic audio and attributing it to its sources (the apps or online services where it is generated), namely synthetic audio attribution. We present a transformer-based model for attributing such synthetic audio to its sources (e.g., mobile applications). We use the PatchOut Fast Spectrogram Transformer(PaSST) for pre-training. We integrate it with a classifier to identify an audio artifact (deciding whether it is natural or generated using a synthetic approach/tool). For the detected fake (i.e., synthetic) audio artifacts, we try to attribute/link to their sources (mobile apps, for example, VoiceAI, ClonyAI, Voice Changer, etc.). Through supervised learning, we have improved the classifier's accuracy in audio classification and helped the understanding of the high-level features that the pre-trained model extracted. We have trained and evaluated natural audio and four mobile application synthesizers. Our experimental results show that our model identifies very accurately and attributes mobile synthesizers with 71% accuracy, using the pre-trained model's feature extraction for complex audio artifacts.

Read the paper · More papers on PaperTik