Enhancing Text-to-SVG Generation via Structured Instruction Embedding and Syntax-Aware Reinforcement in Large Language Models
Peiqing Lu, Shihao Zhao, Y. X. Zhao, Runmian Chang, Yinuo Yang · 2025
Generating Scalable Vector Graphics (SVG) from natural language descriptions poses significant challenges due to the need for precise semantic understanding, structural consistency, and strict syntactic adherence. Existing models often struggle to balance these aspects effectively. This paper proposes SVGGemma-Tuner, a fine-tuning framework that integrates structured instruction embedding to enhance geometric semantic comprehension, a dual-stage decoding architecture to separate layout planning from SVG token generation, and a syntax-aware reinforcement module to optimize syntactic validity through reinforcement learning. By jointly optimizing sequence prediction, spatial alignment, and syntax compliance, SVGGemma-Tuner demonstrates superior performance over existing approaches in generating coherent, semantically accurate, and syntactically valid SVG outputs.