Multilingual Image Captioning Using Blip and Marian MT Models Through a Gradio Interface
Harshitha Dinesh Raghavan, S. Shopika, Dhakshayan Appar, S. Pavithra · 2026
Multilingual image captioning is crucial for connecting visual content to linguistic audiences, but current models lack user-friendly interfaces and high-quality captions. A new framework, BLIP and MarianMT, uses a seamless interface through Gradio to address these limitations. The system ensures accurate captions in English, French, Spanish, and German, making it adaptable for realworld scenarios like accessibility for visually impaired people, educational tools, and automated content creations. The system also overcomes limitations of previous models, including limited language support and inaccuracy in domain-specific contexts. The user-friendly interface makes advanced image captioning accessible to a wider user base.