Comparing PushGP and GPT-4o on Program Synthesis with only Input-Output Examples
Jose Guadalupe Hernandez, Anil Kumar Saini, Gabriel Ketron, Jason H. Moore · Proceedings of the Genetic and Evolutionary Computation Conference Companion · 2025
Genetic programming (GP) and large language models (LLMs) have both achieved notable success in program synthesis. However, the methods for specifying the desired program behavior (i.e., user intent) differ: GP relies on input-output examples, whereas LLMs use text descriptions. In this work, we compare the capabilities of a GP system, PushGP, and an LLM model, GPT-4o, in synthesizing programs where the user intent is specified through input-output examples. Using tasks from the PSB2 program synthesis benchmark, we found that PushGP solved more tasks than GPT-4o. While some tasks were successfully solved by both synthesizers, others were uniquely solved by only one of them, highlighting their complementary strengths. In addition to the prompt with just input-output examples (data-only), we tested GPT-4o with another prompt containing only a textual description of the task (text-only). Both prompt variants successfully solved the same 7 tasks (with different success rates), with the data-only prompt solving an additional task. Ultimately, each synthesizer is successful in distinct ways, highlighting differences in their underlying methodologies.