Automated Captioning for Ergonomic Problem and Solution Identification in Construction Using a Vision-Language Model and Caption Augmentation

Gunwoo Yong, Meiyin Liu, Sang Hyun Lee · 2024

Construction tasks impose high ergonomic risks, making it crucial to observe ergonomic problems (e.g., actions and postures associated with ergonomic risks) and provide solutions. As manually identifying problems and solutions is time-consuming and subjective, there has been extensive development toward automation through computer vision-based or sensor-based applications. Nevertheless, most existing studies have focused on assessing ergonomic risks, leaving tasks of recognizing problems and generating solutions to ergonomists. However, ergonomists are scarce in construction. Therefore, this study aims to automatically identify ergonomic problems and solutions from images by way of image captioning. To overcome limitations of traditional image captioning models, incapable of incorporating knowledge of ergonomics, this study applied a vision-language model (VLM) with caption augmentation leveraging text-based knowledge. The authors tested five work scenarios and showed superior performance of the proposed VLM over the traditional model. This result showed the feasibility of the proposed approach in identifying ergonomic problems and solutions.

Read the paper · More papers on PaperTik