Label-Adversarial Jointly Trained Acoustic Word Embedding

Zhaoqi LI, Li Ta, Qingwei Zhao, Pengyuan ZHANG · IEICE Transactions on Information and Systems · 2022

Query-by-example spoken term detection (QbE-STD) is a task of using speech queries to match utterances, and the acoustic word embedding (AWE) method of generating fixed-length representations for speech segments has shown high performance and efficiency in recent work. We propose an AWE training method using a label-adversarial network to reduce the interference information learned during AWE training. Experiments demonstrate that our method achieves significant improvements on multilingual and zero-resource test sets.

Read the paper · More papers on PaperTik