Enhanced Multilingual Image Captioning: Integrating Text-to-Speech Translation and Sentiment Analysis

Tanuja Konda Reddy, S Veeksha, C. R. Kavitha · 2024

The aim of this project is to create a system that will improve the way we interact with text and images by using the capabilities of computer vision and natural language processing. The main purpose of using this system to generate human-readable image captions is to train a model that can accurately describe an image's contents. The project extends the functionality of the image captioning system by adding language translation inorder to give a multilingual support. The use of a few pre-trained models will enable users with varying language backgrounds use this system to acquire captions in different target languages. This involves integrating language translation techniques into the system and training different models for every target language. Further we extend the project by working in the field of sentiment analysis by identifying the sentiment (positive, negative, or neutral) of captions by examining text and images. For a better understanding of the emotional tone of the information, the system combines CNNs with natural language processing (NLP) algorithms for image analysis and text analysis.

Read the paper · More papers on PaperTik