UniPET-SPK: A Unified Framework for Parameter-Efficient Tuning of Pre-Trained Speech Models for Robust Speaker Verification

Mufan Sang, John H. L. Hansen · IEEE Transactions on Audio Speech and Language Processing · 2025

With excellent generalization ability, self-supervised speech models have shown impressive performance on various downstream speech tasks in the pre-training and fine-tuning paradigm. However, as the size of pre-trained models grows, fine-tuning becomes practically unfeasible due to expanding computation and storage requirements, as well as the risk of overfitting. In this study, we concentrate on exploring parameter-efficient tuning (PET) methods for adapting large-scale pre-trained self-supervised speech models to the speaker verification task. Correspondingly, we propose three parameter-efficient tuning methods: i) an adapter-tuning method, ii) a prompt-tuning method, and iii) a unified framework that effectively incorporates adapter-tuning and prompt-tuning with a dynamically learnable gating mechanism. First, we propose the Inner+Inter Adapter framework, which inserts two types of adapters into pre-trained models, allowing for adaptation of latent features within the intermediate Transformer layers and output embeddings from all Transformer layers, through a parallel adapter design. Second, we propose the Deep Speaker Prompting method that concatenates trainable prompt tokens into the input space of pre-trained models to guide adaptation. Lastly, we propose the UniPET-SPK, a unified framework that effectively incorporates these two alternate PET methods into a single framework with a dynamic trainable gating mechanism. The proposed UniPET-SPK learns to find the optimal mixture of PET methods to match different datasets and scenarios. We conduct a comprehensive set of experiments on several datasets to validate the effectiveness of the proposed parameter-efficient tuning methods. Experimental results on the VoxCeleb, CN-Celeb, and the 1$^{\mathrm{{st}}}$48-UTD forensic datasets demonstrate that the proposed UniPET-SPK can consistently outperform the two PET methods, fine-tuning, and other parameter-efficient tuning methods, achieving superior performance while updating only 5.4% of the parameters. We further conduct experiments on the CN-Celeb and 1$^{\mathrm{{st}}}$48-UTD datasets to demonstrate the robustness and generalization ability of the proposed methods for speaker verification in different languages and more challenging scenarios.

Read the paper · More papers on PaperTik