An Investigation of Context-Driven Caption Generation
Zhishen Yang, Zhishen Yang · Institutional Repositories DataBase (IRDB) · 2025
This thesis investigates context-driven caption generation through systematic studies of two specialized domains: news image and scientific figure captioning.Unlike traditional image captioning, which focuses solely on describing visual content, these two domains require sophisticated integration of contextual information to generate meaningful and accurate captions.The research addresses two fundamental questions: How can models effectively integrate information from visual and textual modalities to generate informative captions, and what is the relative importance of textual versus visual context in context-driven caption generation?The first study on news image captioning demonstrates that generating appropriate captions requires understanding the visual content and its relationship to the broader news narrative.Traditional image captioning approaches are insufficient for this task as they cannot capture the journalistic significance of images within their news context.The study introduces a novel Transformer-based architecture from news articles that effectively integrates visual features with textual context.Through extensive experiments using both automatic metrics and human evaluation, the research reveals that while textual context from news articles provides the primary information for generating contextually appropriate captions, incorporating visual features through the proposed model leads to more context-relevant captions.The model outperforms previous state-of-the-art approaches across multiple evaluation metrics, demonstrating the effectiveness of the transformer-based architecture in handling multimodal information.