Computational Integration of Human Vision and Natural Language through Bitext Alignment

Preethi Vaidyanathan, Emily Tucker Prud’hommeaux, Cecilia Ovesdotter Alm, Jeff B. Pelz, Anne Reeves Haake · 2015

Multimodal integration of visual and linguistic data is a longstanding but crucial challenge for modeling human understanding.We propose a framework that uses an unsupervised bitext alignment method to integrate visual and linguistic data.We present an empirical study of the various parameters of the framework.Our results exceed baselines using both exact and delayed temporal correspondence.The resulting alignments can be used for image classification and retrieval.

Read the paper · More papers on PaperTik