Image Caption Generation using CNN and Transformers

Konda Rupa Narendra Babu, Muchhu Lakshmi Poojitha, Nagisetti Jyothi, Yandrapati Gnana Kishor, K Vinay Chandrasekhar · 2025

Image caption generation from the combination of computer vision with NLP is a critically important task for machines being able to describe images, and this project leverages the power of CNN architectures and Transformer architectures to do precisely that: generate context-aware captions. CNN Encoder: EfficientNetB0 - The CNN encoder is applied to extract high-level features from images. Then a Transformer encoder-decoder model is used to generate word-by-word captions. The relationships between image features are captured, and semantically rich, syntactically correct descriptions are produced. This approach yields an efficient and scalable solution to automatic image captioning with potential applications in accessibility, search engines, and social media and demonstrates the strength of deep learning techniques for this task.

Read the paper · More papers on PaperTik