Cross-lingual XLSR-Wav2Vec2-based Speech Spoofing Detection for Vietnamese Speech

Thi-Thanh-Hang Nguyen, Thi-Ngoc-Diep Do · 2025

The Automatic Speaker Verification (ASV) via human speech is one of the most popular recognition methods which has been recently applied in many security control systems. The spoofing speech detector is therefore usually integrated into the ASV system to make the system more robust to spoofing attacks in which the attackers mimic voices belonging to a specific individual. We focus on the spoofing speech detection for Vietnamese speech. Due to the lack of ASV data for Vietnamese, a spoofing speech detector based on cross-lingual speech representation is proposed which can leverage the big speech data from high-resource language. The proposed XLSR-Wav2Vec2-based Spoofing Speech Detector model uses the XLSR-Wav2Vec2 cross-lingual self-supervised speech representation model as the encoder and a classification network as the backend of the model to classify the spoof and bona fide speech. Three evaluations have been performed: (1) fine-tuning the model on the ASVspoof2019 dataset and evaluating it on the ASVspoof2021 dataset to assess generalization, (2) testing the English-adapted model directly on Vietnamese data to evaluate cross-lingual transfer, and (3) further fine-tuning the model on the Vietnamese corpus to enhance its performance on the target language, with comparisons to a RawNet2 baseline trained directly on the Vietnamese data. The results show that the proposed method effectively learned crucial features for discriminating between spoof and bona fide speech. For the first test, the EER and accuracy metrics on English test set were 4.34 and 98.9% respectively which are comparable to the average result of the current systems on English. And even when fine-tuning and testing on two different languages, EER and accuracy metrics of the second test were 14.15 and 81.87%. The last test on fine-tuning helped to improve the quality of SSD on Vietnamese speech to 1.41 for EER and 97.10% for accuracy metrics. The results of these evaluations indicate the speech representation can be shared in this cross-lingual model to improve the quality of the spoofing speech detector. The Vietnamese dataset has been publicly released on Kaggle to facilitate reproducibility and future research.

Read the paper · More papers on PaperTik