Zooming in on Zero-Shot Intent-Guided and Grounded Document Generation using LLMs
Pritika Ramu, Pranshu Gaur, Rishita Emandi, Himanshu Maheshwari, Danish Javed, Aparna Garimella · 2024
Repurposing existing content on-the-fly to suit author's goals for creating initial drafts is crucial for document creation.We introduce the task of intent-guided and grounded document generation: given a user-specified intent (e.g., section title) and a few reference documents, the goal is to generate section-level multimodal documents spanning text and images, grounded on the given references, in a zero-shot setting.We present a data curation strategy to obtain general-domain samples from Wikipedia, and collect 1,000 Wikipedia sections consisting of textual and image content along with appropriate intent specifications and references.We propose a simple yet effective planningbased prompting strategy Multimodal Plan-And-Write (MM-PAW), to prompt LLMs to generate an intermediate plan with text and image descriptions, to guide the subsequent generation.We compare the performances of MM-PAW and a text-only variant of it with those of zero-shot Chain-of-Thought (CoT) using recent close and open-domain LLMs.Both of them lead to significantly better performances in terms of content relevance, structure, and groundedness to the references, more so in the smaller models (upto 12.5 points ↑ in Rouge 1-F1) than in the larger ones (upto 4 points ↑ R1-F1).They are particularly effective in improving relatively smaller models' performances, to be on par or higher than those of their larger counterparts for this task.