Multimodal NLP for image captioning : Fusing text and image modalities for accurate and informative descriptions
Manisha Tiwari, Pragati Khare, Ishani Saha, Mahesh Mali · Journal of Information and Optimization Sciences · 2024
Multimodal Natural Language Processing (NLP) offers significant potential for improving the understanding and generation of content that combines various modalities, including text and images. Image captioning, automatically generating textual descriptions of images, represents a crucial application of multimodal NLP. While existing methods primarily rely on image features, we propose a novel multimodal NLP model that leverages the power of both text and image modalities for generating informative and accurate image captions. Our model incorporates information from text descriptions associated with images, enabling it to capture contextual cues and generate richer captions than traditional unimodal models. We validate our method using the industry standard Flickr8K dataset and obtain cutting-edge outcomes, proving the potency of our multimodal fusion technique. Furthermore, we discuss the challenges and opportunities of multimodal NLP for image captioning and highlight its potential to revolutionise how we interact with computers and interpret visual information.