An Empirical Study of Security Risks for the Web Code Generation by ChatGPT

Mingrui Hu · Applied and Computational Engineering · 2025

Large language models (LLMs) have demonstrated remarkable capabilities in code generation and semantic understanding, enabling ordinary users to generate their own software systems using natural language instructions. This study takes website systems as a case to investigate a user-centered paradigm for code generation and its evaluation. First, users submit their requirements to the LLM via a web interface, prompting the model to automatically generate website project code. Then, through a set of prompt engineering methods and quantitative evaluation techniques developed for this study, we conduct a multi-dimensional assessment of the quality and security of the generated website systems using different types of LLMs and varying system function weights. A hybrid evaluation strategy is proposed to integrate and optimize assessment results across different LLMs. Evaluation dimensions include the degree to which user requirements are satisfied, completeness of website functionality, potential security risks, and code reliability. This research introduces evaluation criteria such as automated review models, functional coverage, and static vulnerability analysis to explore the feasibility, advantages, and limitations of using LLMs as both code generators and reviewers. The findings contribute to our understanding of the practical value of multi-agent LLM collaboration in software development and reveal major current challenges such as functional hallucination, incomplete implementation, and overly optimistic evaluation mechanisms.

Read the paper · More papers on PaperTik