Procedural Text Generation from a Photo Sequence

Taichi Nishimura, Atsushi Hashimoto, Shinsuke Mori · 2019

Multimedia procedural texts, such as instructions and manuals with pictures, support people to share how-to knowledge.In this paper, we propose a method for generating a procedural text given a photo sequence allowing users to obtain a multimedia procedural text.We propose a single embedding space both for image and text enabling to interconnect them and to select appropriate words to describe a photo.We implemented our method and tested it on cooking instructions, i.e., recipes.Various experimental results showed that our method outperforms standard baselines.

Read the paper · More papers on PaperTik