A Hybrid CNN and LSTM for Image Caption Generation
Uma Rani V, K. Revathy · 2024
Image captioning, a complex job in artificial intelligence, has received a lot of attention for its numerous natural language processing applications. Because of its inherent layering and encoder-decoder architecture, the deep learning model is the most commonly employed for picture captioning. In this study, a CNN-LSTM model is used to create visual descriptions. This approach uses Flickr8K (8,000 photos) for training the models, which helps to reach 94% accuracy. This image captioning system was trained using an 8,000-images from Flickr8k. The experimental result for the hybrid model is 94% accuracy, which is greater than the CNN model.