WebCoder: Multimodal Approach for Automated Web Service UI Code Generation

Huafeng Su, Lei Yu, Conghui Yang · 2025

In recent years, the remarkable success of large language models has substantially advanced research in code generation. However, most current research on code generation predominantly focuses on unimodal approaches, while studies that explore the integration of images for code generation remain remarkably scarce. This limitation is particularly pronounced in the context of generating front-end code from Web service UI designs, a task that poses numerous challenges. Firstly, existing datasets are both insufficient and of low quality, often being synthesized through generative models. Secondly, current Web UI code generation models not only struggle to produce high-quality code but also lack the capability to generate interactive and dynamic code. To address these issues, we constructed two high-quality datasets specifically for the task of generating front-end code from Web service UI designs. The datasets contain a rich collection of web screenshots, the corresponding front-end code, and related instructions. This paper presents an innovative multimodal approach that integrates multi-stage fine-tuning and multi-step prompting strategies to generate Web front-end code (HTML, CSS, JavaScript). This approach substantially enhances the efficiency and accuracy of the code generation process, mitigates errors in the generated code, and improves the consistency of web page restoration. Through extensive experimental validation, our model has demonstrated strong competitiveness on multiple public benchmarks, achieving or surpassing GPT-4V on several metrics, with performance levels above 90 %, highlighting the significant potential of our approach in practical applications.

Read the paper · More papers on PaperTik