Word Discovery in Visually Grounded, Self-Supervised Speech Models

Puyuan Peng, David F. Harwath · Interspeech 2022 · 2022

OursFigure 1: HuBERT: sum of attention weights each frame receives from other frames.Ours (VG-HuBERT3): attention weights each frame receives from the [CLS A] token.Attention weights from different attention heads are coded with different colors.

Read the paper · More papers on PaperTik