Dense Image Captioning in Hindi

K. Gill, Sriparna Saha, Santosh Kumar Mishra · 2021 IEEE International Conference on Systems, Man, and Cybernetics (SMC) · 2021

Much work has been done on dense image captioning in the English language. In this paper, we propose a Encoder-Decoder architecture for the first time, to the best of our knowledge, to generate Hindi dense image captions. We leverage the proven encoder-decoder architecture and provide a Hindi data-set along with results of dense captioning on the data-set. We also propose a novel post-processing technique to further improve the quality of captions by using multiple discriminators (language discriminator, length penalty and dissimilarity score). Hindi is the fourth most spoken language in the world. It is the official language of India and to the best of our knowledge, this is the first attempt at dense image captioning in the Hindi language. The data-set is manually created by translating Visual Genome captions from English to Hindi under manual supervision. Both qualitative (in terms of adequacy and fluency) and quantitative (in terms of BLEU scores) analysis of the generated captions reveal the effectiveness of our proposed approach compared to several strong baselines.

Read the paper · More papers on PaperTik