Simplifying Image Captioning in Hindi with Deep Learning
Shimpi Rawat, Manika Manwal, Kamlesh Chandra Purohit · 2024
We aim to simplify and optimize the process of generating Hindi captions for visual content using deep learning models. With the advancements in deep learning, generating textual descriptions for images is now possible. Any spelling, grammar, and punctuation errors have been corrected. However, extending this capability to languages other than English presents notable challenges. To bridge this gap, we propose a streamlined approach for generating image captions in Hindi. Our proposed methodology utilizes a neural network model that employs a multilayered CNN-LSTM architecture with attention mechanisms to enhance caption generation. The model has been trained on a dataset that has been sourced from MSCOCO and we have adapted and optimized hyper parameters to increase the probability of generating accurate Hindi descriptions. To evaluate the effectiveness of our models, we use BLEU scores that demonstrate significant improvements compared to existing work in this domain. Our research holds great significance beyond just technological innovation. We acknowledge the potential impact of automatic description of events in the environment and their translation into captions or messages on society. By simplifying the image captioning process in Hindi, we aim to promote linguistic inclusivity and provide a convenient interface for Hindi speakers, encouraging them to engage more easily with the diverse digital landscape. As we embark on this journey, our dedication goes beyond the intricacies of deep learning. Our goal is to simplify and make the process of generating image captions accessible and effective for the Hindi-speaking community. This research is in line with the wider discussion on linguistic diversity, cultural relevance, and the democratization of digital experiences in the ever-changing landscape of language technology.