Contextual Caption Generation Using Attribute Model
Jessin Donnyson, Masayu Leylia Khodra · 2020
Image captioning has been a big field of research and gained many enterprises' attention in the past few years. Unfortunately, there are still many issues preventing the general usage of image captioning. One of the pitfalls is lack of context. The common approach to solve this issue is to retrain the model with a suitable dataset for every purpose, but this approach is expensive and time-consuming. We propose a method to shift words' distribution and guide the captioning process towards the desired result using a gradient vector acquired from attribute models. This method requires neither retraining captioning models nor customized datasets. In conclusion, this approach shows positive results even though the usage is still limited and highly dependent on the added context and parameter used.