Image to Text Conversion: State of the Art and Extended Work

Nada Farhani, Naim Terbeh, Mounir Zrigui · 2017

The aim of this article is to study the conversion of information between the different modalities (text, image) due to the evolution of human-machine communication that introduced the use of natural communication modalities to humans such as gestures, speech, sound and vision. In fact, one of the main challenges of this "multimodal" learning is the learning of a shared representation between the distinct modalities and the prediction of the missing data (for example, by retrieval or synthesis) from a conditioned modality to another. Some researches work on the different types of conversions; Text to Speech, Speech to Picture or Text to Picture synthesis and vice-versa but in this paper we will focus on: Text to Picture (TTP) and Picture to Text (PTT) synthesis.

Read the paper · More papers on PaperTik