A Probabilistic Model for Joint Learning of Word Embeddings from Texts and Images
Melissa Ailem, Bowen Zhang, Aurélien Bellet, Pascal Denis, Fei Sha · 2018
Several recent studies have shown the benefits of combining language and perception to infer word embeddings.These multimodal approaches either simply combine pre-trained textual and visual representations (e.g.features extracted from convolutional neural networks), or use the latter to bias the learning of textual word embeddings.In this work, we propose a novel probabilistic model to formalize how linguistic and perceptual inputs can work in concert to explain the observed word-context pairs in a text corpus.Our approach learns textual and visual representations jointly: latent visual factors couple together a skip-gram model for co-occurrence in linguistic data and a generative latent variable model for visual data.Extensive experimental studies validate the proposed model.Concretely, on the tasks of assessing pairwise word similarity and image/caption retrieval, our approach attains equally competitive or stronger results when compared to other state-of-the-art multimodal models.