WaveTrap: A Clean-Label Poisoning Attack Against Unauthorized Voice Data Collection

Yuanjie Zhang, Yunjie Ge, Lingchen Zhao, Qian Wang · IEEE Network · 2025

Voice data is essential for training audio intelligence systems, yet its collection often faces ethical and logistical barriers due to privacy concerns. The high cost of acquiring legally consented voice data incentivizes malicious actors to covertly harvest sensitive acoustic information. To counteract such threats, we propose Wave-Trap, a clean-label poisoning attack designed to sabotage voice data while preserving its human-perceivable intelligibility. Our method introduces imperceptible perturbations into audio signals, manipulating their Mel Frequency Cepstral Coefficients (MFCCs)—a critical feature representation in audio processing—to embed stealthy poisoning patterns. When poisoned data is utilized for model fine-tuning or training, these perturbed MFCCs degrade the efficacy of downstream audio intelligence tasks. Extensive experiments demonstrate the efficacy of WaveTrap against two typical audio recognition systems: speaker recognition (SR) and speech command recognition (SCR). For instance, the accuracy of VggVox drops by 30.34% when only 10% of the poisoned voice data is used for fine-tuning, and by 7.28% when 10% of the training data is poisoned.

Read the paper · More papers on PaperTik