Capturing Semantic Knowledge In Object Localization In Captioning Images
Chaitrali Prasanna Chaudhari, Satish Devane · 2021
Image captioning is the process of generating the description of an image in a human-understandable form that needs to be meaningful, syntactically, and semantically correct. This process of generating description by understanding the semantic knowledge in the `scene' is very natural for humans since they can easily extract the common meaning from the perceived image focusing on salient objects and rejecting unimportant aspects and express the `story' behind it in meaningful words. The same task is much complex for machines because the semantic knowledge or context can be evenly or unevenly spread in an image as well as it can be local or global. To achieve the purpose of meaningful image caption generation there is a need of combining the research communities of natural language processing and computer vision. This paper presents a survey of various approaches used in image captioning to improve the context in description generation by capturing the semantics.