ENVISAGE: An image captioning dataset and conformal uncertainty modeling framework to assist visually impaired individuals
Özkan ÇAYLI, Volkan Kılıç, Wenwu Wang · Pattern Recognition · 2026
Image captioning has become a key task in vision-language research, bridging visual understanding and natural language generation. However, existing image captioning datasets and models are primarily optimized for sighted individuals, often failing to address the needs of visually impaired individuals who rely on captions for scene understanding and navigation. To address this gap, we introduce ENVISAGE, a large-scale image captioning dataset curated with input from visually impaired individuals. It contains 53,452 images and 267,260 accessibility-focused captions that emphasize spatial relations and object affordances. We benchmark both baseline ( Show and Tell , Show, Attend and Tell ) and state-of-the-art ( BLIP and BLIP-2 ) captioning models on ENVISAGE under zero-shot and fine-tuned settings. Beyond the dataset contribution, we present the first framework for conformal uncertainty estimation in image captioning, providing finite-sample coverage guarantees that improve confidence in caption quality. Results demonstrate that conformal predictions achieve reliable coverage guarantees across diverse image categories. Among all evaluated models, the fine-tuned BLIP-2 OPT6.7B achieves the best overall performance, with a final score of 0.4362 across the image captioning metrics. Furthermore, uncertainty-aware evaluation shows that captions in the low-uncertainty band attain the highest quality, achieving a final score of 0.4476, confirming that the proposed conformal framework produces calibrated uncertainty estimates that align with caption reliability. ENVISAGE and our uncertainty-aware framework establish a robust and accessible foundation for image captioning systems.