Enhanced image understanding via deep learning
R.S. Latha, C. Roopa, D. Dhavamani, E.K. Dayanithi, A.K. Arumugam, R. Dhanushkumar · 2026
Image captioning is the process of generating textual descriptions of an image, using a combination of natural language processing for language modeling and computer vision for image comprehension. Although considerable research has been done in English, not much work has been performed in regional languages like Tamil. This paper represents the first known attempt in this field and presents a novel method for annotating Tamil images using the Flickr30k dataset. A dataset is created by taking the Flickr30k English captions and manually translating into Tamil. The proposed uses CNNs based on the VGG16 model together with LSTMs which model the sequence to generate grammatical Tamil captions. This new translated dataset is used to train the model, which splits into training, validation, and test subsets, and evaluates results using BLEU scores. The model showed encouraging results in both accuracy and fluency. Apart from improving picture captioning for regional languages, this work develops assistive technologies for people speaking Tamil and paves the door for further studies in low-resource languages.