Pitch-Shifted Speech Detection using Audio SSL Model
Artem Chirkovskiy, Marina Volkova, Nikita Khmelev · 2025
Pitch-shifting is a common yet effective method for speech processing, often exploited by phone fraudsters to conceal their identities. In this work, we evaluate the performance of pitch-shifted speech detection using a self-supervised learning (SSL) model, WavLM, broadly used for deepfake detection. The LibriTTS dataset serves as the source of genuine speech, while spoofed speech is generated using ten pitch-shifting methods, including TD-PSOLA, WSOLA, OLA, and the Phase Vocoder. To simulate real-world phone fraud scenarios, we apply the G.711 and G.729 telephony codecs to assess detection robustness in degraded audio conditions. Additionally, we analyze the anonymization capabilities of different pitch-shifting techniques using an open-source Automatic Speaker Verification (ASV) system. Our results reveal two key findings: (1) artifacts introduced by pitch-shifting methods are distinct and do not generalize well across techniques, and (2) subtle pitch changes (±1 semitone) preserve speech naturalness but fail to anonymize speakers effectively. Nevertheless, larger pitch shifts become harder to detect under unseen codec conditions, highlighting a critical trade-off between detectability and anonymization.