Learning Image Embeddings using Convolutional Neural Networks for Improved Multi-Modal Semantics

Douwe Kiela, Léon Bottou · 2014

We construct multi-modal concept representations by concatenating a skip-gram linguistic representation vector with a visual concept representation vector computed using the feature extraction layers of a deep convolutional neural network (CNN) trained on a large labeled object recognition dataset.This transfer learning approach brings a clear performance gain over features based on the traditional bag-of-visual-word approach.Experimental results are reported on the WordSim353 and MEN semantic relatedness evaluation tasks.We use visual features computed using either ImageNet or ESP Game images.

Read the paper · More papers on PaperTik