Potential of Incorporating Motion Estimation for Image Captioning
Kiyohiko Iwamura, Jun Younes Louhi Kasahara, Alessandro Moro, Atsushi Yamashita, Hajime Asama · 2021
Automatic image captioning has various important applications such as indexing images on the Web or the depiction of visual contents for the visually impaired. Recently, deep learning based probabilistic frameworks have been greatly researched for image captioning. However the existing deep learning methods are only established on visual features, which have problems generating captions related to motions, because visual features from images do not include motion features. In this paper, we propose a novel, end-to-end trainable, deep learning image captioning model that estimates motion features from a image to help generate captions. Our proposed model was evaluated on two datasets, MSR-VTT2016-Image, and several copyright free images. We demonstrate that our proposed method using motion features improves performance on caption generation and that the quality of motion features is important to generate captions.