Memory Storable Network Based Feature Aggregation for Speaker Representation Learning
Bin Gu, Wu Guo, Jie Zhang · IEEE/ACM Transactions on Audio Speech and Language Processing · 2023
Learning fixed-dimensional speaker representation using deep neural networks is a key step in speaker verification. In this work, we propose an auxiliary memory storable network (MSN) to assist a backbone network for learning discriminative features, which are sequentially aggregated from lower to deeper layers of the backbone. The proposed MSN has a similar architecture to the ResNet and contains a set of cascaded feature aggregation (FA) blocks. Each FA block first aggregates the multi-level features from the previous block and the features from the corresponding backbone layer. The output features of each intermediate layer within the backbone are then refined by the multi-level features of the corresponding FA block through masking and biasing operations. Finally, the features from the last layers of both MSN and the backbone are concatenated to form more discriminative speaker representations. Experimental results on five public datasets show significant and consistent improvements over conventional approaches. The effectiveness of the proposed method is also validated using ablation studies, showing a robust generalization capacity in combination with different backbone networks.