Show, Prefer and Tell: Incorporating User Preferences into Image Captioning
Annika Lindh, Robert Ross, John D. Kelleher · 2023
Current image captioning models produce fluent captions, but they rely on a one-size-fits-all approach that does not take into account the preferences of individual end-users. We present a method to generate descriptions with an adjustable amount of content that can be set at inference-time, thus providing a step toward a more user-centered approach to image captioning.*