Learning Words from Images and Speech

Gabriel Synnaeve, Verteegh Maarten, Emmanuel Dupoux · Figshare · 2014

This paper explores the possibility to learn a semantically-relevant lexicon from images and speech only. For this, we train a multi-modal neural network working both on image fragments and on speech features, by learning an embedding in which images and content words that co-occur together are close. Making no assumption on the acoustic model, this paper shows promising results on how multi-modality could help word learning.

Read the paper · More papers on PaperTik