A Dynamic Convolution Framework for Session-Independent Speaker Embedding Learning
Bin Gu, Jie Zhang, Wu Guo · IEEE/ACM Transactions on Audio Speech and Language Processing · 2023
Speaker verification (SV) has suffered from session variability in complex acoustic scenarios, and learning session independent speaker representations remains a challenging problem. To tackle this, we propose a dynamic convolution framework for SV in this article, which dynamically adapts the model parameters to each input feature during inference, such that the model can flexibly extract robust speaker characteristics under different acoustic conditions. Specifically, we combine an adaptive context vector extraction (ACVE) module and a sub-kernel scaling (SKS) module in a cascaded manner. The ACVE uses a moving weighted sum and a sub-band self-attention in parallel to extract global-local information, which is then passed through the subsequent SKS module for efficient dynamic kernel generation. The proposed method can be easily implemented on the backbone network by replacing the conventional static counterparts. Experimental results on five public SV datasets show significant and consistent improvements over comparison approaches, and ablation studies and visual analysis further demonstrate the effectiveness of the proposed method.