Multi-Task Diffusion Model for Simultaneous Text and Image Inpainting

Bolin Wang, Xinyi Shen, Kejun Zhang · 2024

Text image inpainting is critically important across a wide range of applications, from archival recovery to image editing. This task presents significant challenges, as it requires preserving text legibility while maintaining overall image integrity. In this paper, we tackle the problem of inpainting corrupted text images by proposing a novel framework that utilizes multi-task learning within a diffusion model. Unlike traditional two-stage approaches, which first restore the global text structure and then the image content, our method simultaneously repairs both the text structure and the image content, enabling effective information sharing between the two tasks. This results in a more robust and generalized model capable of handling diverse corruption scenarios more effectively. We validate the effectiveness of our approach through extensive experiments and comparisons with existing methods, demonstrating notable improvements in both image quality and accuracy in downstream tasks.

Read the paper · More papers on PaperTik