Dimension-Specific Margins and Element-Wise Gradient Scaling for Enhanced Matryoshka Speaker Embedding

Sunchan Park, Hyung Soon Kim · IEEE Access · 2025

Speaker embeddings trained with Matryoshka Representation Learning (MRL) provide embeddings of various dimensions with minimal overhead, adapting to different computational and storage constraints. Compared to single-dimensional models, MRL-based models show improved speaker verification performance in lower dimensions, but there is some degradation in higher dimensions. Analyzing learned embeddings, we observe an imbalance in the element-wise magnitudes of speaker embeddings trained with MRL. Specifically, the higher-dimensional elements exhibit extremely small values, which could reduce their contribution to cosine similarity and degrade speaker verification performance. To address this imbalance and improve performance consistency across all dimensions, we propose two methods: dimension-specific margins and element-wise gradient scaling. Dimension-specific margins stabilize training by adjusting the margin for each dimension to mitigate instability caused by excessively high values for given dimensions. Element-wise gradient scaling mitigates imbalance by scaling gradients propagated to each element, considering differences in embedding dimensionality and the number of loss functions influencing each element. Evaluation on the VoxCeleb benchmark shows that the proposed methods, when applied to MRL, improve speaker verification performance in higher-dimensional embeddings while maintaining performance in lower-dimensional embeddings. Additionally, an analysis of the element-wise magnitudes of learned embeddings visually demonstrates that element-wise gradient scaling effectively mitigates the magnitude imbalance.

Read the paper · More papers on PaperTik