Echoes Unveiled: Identifying Synthetic Voices

Daniel Pluth, Jordan Hosier, Yu Zhou, Vijay K. Gurbani · 2025

The advent of deep neural networks in natural language and speech processing has created a new attack vector in the form of synthetic voices cloned from less than 30-seconds of voice sample from a human counterpart. Effectively detecting spoofing attacks is critical for any speech application that uses voice for authentication, verification, and identification. With the rapid rise of highly effective speech synthesis, it is challenging to identify synthetic voices while generalizing to novel voices, synthesizers, and channel conditions. In this paper, we present a model aimed at identifying synthetic voices and demonstrate its effectiveness and generalizability. We further motivate the need for the research community to consider channel conditions when detecting voice spoofing. Our work demonstrates that channel conditions play an inordinate role in identifying a spoofed voice, and detection techniques that do not consider variable channel conditions will exhibit high error rates.

Read the paper · More papers on PaperTik