CaptionPS: Benchmarking and Analyzing Worker-Centric Detailed Image Captioning in Power Scenarios

Zixiang Wang, Xin Chen, Jiajia Han, Junyu Cai, S Wang, Junjie Huang · IEEE Access · 2026

Image captioning has long been a persistent challenge in the field of vision-language research. With the rise of Large Language Models (LLMs), modern Vision-Language Models (VLMs) can produce effective captions in general-purpose scenarios. However, when confronted with complex power safety violation scenarios, they often omit critical violation details, greatly limiting their reliability in safety supervision. Therefore, we propose CaptionPS, an image-captioning dataset tailored for violation detection. CaptionPS is constructed through a systematic pipeline with multiple rounds of manual annotation. Each sample includes not only the classification of the violation scene and the description of the image content, but also the prediction of the safety of subsequent behaviors, ensuring the safety of workers. Based on CaptionPS, We evaluate a suite of VLMs. Experimental results show that open-source pretrained models generally exhibit low performance. However, after supervised fine-tuning, the overall performance of all models has been significantly improved. Among the evaluated models, InternVL3.5-8B achieves the highest performance on detailed captioning. Meanwhile, InternVL3.5-8B and Llama-3.2-11B achieve the best performance on scene classification and safety prediction tasks, respectively, with accuracy rates reaching 87.32% and 78.51%, which are 48.47% and 45.66% higher than those of the baseline models. The proposed CaptionPS helps MLLMs better learn violation characteristics and improve scene recognition in power scenarios.

Read the paper · More papers on PaperTik