Multi-Task Diffusion Model for Simultaneous Text and Image Inpainting
Bolin Wang, Xinyi Shen, Kejun Zhang · 2024
Text image inpainting is critically important across a wide range of applications, from archival recovery to image editing. This task presents significant challenges, as it requires preserving text legibility while maintaining overall image integrity. In this paper, we tackle the problem of inpainting corrupted text images by proposing a novel framework that utilizes multi-task learning within a diffusion model. Unlike traditional two-stage approaches, which first restore the global text structure and then the image content, our method simultaneously repairs both the text structure and the image content, enabling effective information sharing between the two tasks. This results in a more robust and generalized model capable of handling diverse corruption scenarios more effectively. We validate the effectiveness of our approach through extensive experiments and comparisons with existing methods, demonstrating notable improvements in both image quality and accuracy in downstream tasks.