Fontstyler: Artistic Font via Two-Stage Stylized Diffusion

Xu-Yao Zhang, Yu Feng Wu · 2025

Generating artistic fonts is significantly challenging because it requires a subtle balance between readability and artistic degree while presenting the visual expression of a specific subject in pretty and rational changes in the shapes and textures of target letters. While existing studies integrate the textures from subject reference images into target letters, the artistic degree and diversity of outcomes are constrained. By adopting Latent Diffusion Model (LDM), DS-Fusion generates more expressive and plentiful artistic fonts. Nonetheless, font structure fractures within DS-Fusion results. Furthermore, DS-Fusion may not sufficiently blend the semantics of the style prompt into the results. These issues undermine the aesthetic appeal and legibility of artistic fonts. To alleviate these issues, we introduce Fontstyler, an extension of DS-Fusion, employing a two-stage training process to blend style prompt semantics into target letters further and alleviate font structure fractures within DS-Fusion outputs. Specifically, we make LDM to blend prompt semantics into target letters by employing a CNN-based discriminator in the first training process. This discriminator discerns between the target letter image (considered real samples) and the outputs of reverse diffusion in LDM (considered fake samples). In the second training process, we take the results of the first training process as inputs to alleviate the font structure problem and improve the artistic degree of our model's results. We demonstrate the quality of our method on numerous examples and compare it to various baselines.

Read the paper · More papers on PaperTik