Advance Image Captioning Using Machine Learning Techniques
Mayank Sisodia, Dhruv Varshney, Ajendra K. Sharma, K. Suresh · 2023
This study proposes a model for picture captioning that uses both convolutional neural networks (CNN) architecture and long-short-term memory (LSTM). This can serve the use in generating the descriptions of the visual content present in e-learning platforms or online tutorials, so for the better comprehension and the support of the individuals. Also, our model can serve its uses in social-media platforms for better experience. The suggested model employs a pre-trained CNN Dense Net architecture to extract characteristics from the input image before utilizing an LSTM to provide a descriptive caption. These extracted characteristics are inserted into an LSTM network to produce descriptive captions. To improve the accuracy and fluency of the generated captions, we also incorporate a semantic attention mechanism that allows the model in order to emphasize on key aspects of the image during the captioning process. We test our model on a benchmark dataset, Flickr8k, and assess its performance compared to other cutting-edge methods. The experimental findings show that the suggested model performs competitively on both datasets and outperforms current models regarding caption quality and diversity.