Is That Me? Using Speaker Identity to Detect Fake Speech

Shilpa Chandra, Padmanabhan Rajan · 2024

Of late, techniques to generate fake speech have become more and more sophisticated, resulting in challenges in their detection. This paper explores using speaker information in detecting fake speech. Speaker information derived from the linear prediction residual signal is used to supplement a state-of-the-art fake speech detector. Multi-branch convolutions followed by a transformer encoder is used to represent the residual signal compactly. Further, by novel utilization of a contrastive loss function, speaker information is captured effectively from a given utterance and an additional genuine utterance from the same speaker. On evaluation under a speaker-aware protocol, the proposed method shows promise in fake speech detection accuracy on the ASVspoof 2019 and ASVspoof 2021 datasets. Additionally, several ablation studies reveal the effectiveness of the residual signal in capturing speaker information. Code is available at https://github.com/shilpac131/SADD/

Read the paper · More papers on PaperTik