One-Shot Font Generation with Masked Diffusion Transformers
Xijia Wang, Yefei Wang, Ai Chen, Jinshan Zeng · 2024
Font generation is aims to create a new font library by synthesizing characters with specific style references. Existing font generation methods can generally be categorized into two primary types: those based on GANs and those utilizing diffusion models. Owing to adversarial training, GAN-based approaches often encounter issues with training instability and inaccurate character generation. Recently, in the font generation task, the diffusion model has shown outstanding performance due to the stability of its training process. However, existing methods primarily rely on U-Net architecture, and the effectiveness of Transformer architecture has not been fully explored in the field of font generation. To address this, we introduce a Diffusion Transformer-based approach for font generation. Besides, we incorporate masked modeling into the font generation task, allowing the model to better grasp the overall integrity of each character. Extensive experiments demonstrate that our approach yields superior generation results compared to mainstream approaches. Additionally, cross-language font generation experiments confirm our method's effectiveness in generating fonts across diverse languages.