VoiceCap: Automatic Image Captioning and Voice Generation
Sankari Subbiah, G. Sudha, S. Saranya, S Bharathan · 2025
One of the most in-demand resources in the contemporary period is picture captioning. In addition, there are built-in apps that use models from deep neural networks to create and provide a caption for a specific picture. picture captioning refers to the process of creating a description for a picture. It entails identifying key elements in a picture, as well as those items properties and the connections between them. Syntactically and semantically sound sentences are produced by it. VoiceCap is an innovative software for smartphones that uses sophisticated deep learning algorithms to automatically provide audio descriptions for photos and videos. Making sure people with all kinds of visual impairments can enj oy visual material to the fullest is its primary goal in advocating accessibility and inclusion. The VoiceCap cross-platform solution uses a combination of computer vision and natural language processing algorithms to generate spoken and text-based captions for user-uploaded photographs and videos. The captions' correctness and relevance to context are guaranteed by a state-of-the-art Vision-Encoder-Decoder Model.