A Two-Stage Code Generation Method using Large Language Models

Zhao Dapeng, Geng Tongcheng · International Journal of Performability Engineering · 2024

Large language models are capable of generating source code in a zero-shot manner to develop programs that meet user functional requirements.However, when faced with scenarios involving complex business requirements, the generated source code may fail to satisfy user needs.Addressing the challenge of understanding software requirements, we propose a two-stage code generation approach.Initially, the large language model generates pseudocode based on the user's functional requirements, refined through an iterative process with user feedback.Subsequently, the model generates source code based on the finalized pseudocode.We conducted empirical studies on an open code generation dataset, and experimental results with models such as GPT-4, Claude Sonnect 3, and Geminipro 1.5 demonstrate that our method outperforms zero-shot prompt learning in scenarios with complex user requirements, with improvements in PASS@K reaching up to 15%.

Read the paper · More papers on PaperTik