Enhancing Automatic Image Captioning: A Systematic Analysis of Models, Datasets, and Evaluation Metrics
Kothai Ganesan, S. Amutha, G. Venkata Saketh, G. Vishnu Vardhan Reddy, G. Harinath Chowdary, J. Shiva Sai · 2024
The fusion of computer vision and natural language processing (NLP) has given rise to the interdisciplinary field of automatic image captioning, which aims to generate descriptive text for images without human intervention. This area of research has attracted considerable interest due to its potential uses in improving accessibility, managing content, and enhancing search engine optimization. Over the past years, there has been a drastic increase in the accuracy and detail of captions thanks to recent advances in deep learning, especially with convolutional neural networks (CNNs), long short-term memory (LSTMs), and recurrent neural networks (RNNs). This research examines cutting-edge approaches to automatic image captioning, including attention mechanisms, models based on transformers, and techniques utilizing reinforcement learning. The effectiveness of these methods was evaluated using benchmark datasets, and their advantages and limitations were assessed. Additionally, the obstacles faced in producing contextually appropriate and linguistically precise captions were explored, such as dealing with intricate scenes, uncommon objects, and abstract ideas. The paper concludes by highlighting potential areas for future research, including learning across multiple modalities, captioning across languages, and incorporating common-sense knowledge to improve caption generation.