Prompt Optimization with Human Annotations Using Pairwise Comparison

Yung-En Lu, Shin-Jie Lee · 2024

The performance of Large Language Models (LLMs) is highly dependent on the input prompt, which has led to several recent studies focused on prompt optimization. Optimizing prompts for generation tasks that align with user intent is particularly challenging due to the subjective nature of these tasks and the prohibitive costs associated with the extensive human annotations required for repeated evaluations. This paper addresses the problem of optimizing prompts for LLMs in generation tasks by utilizing user pairwise comparison feedback instead of numerical scores. We propose a framework that integrates user pairwise comparison feedback into the prompt optimization process by maintaining a dynamically ranked set of prompts. In each iteration, a new prompt is generated based on previous user pairwise annotations. We evaluated our approach and an existing score-based method on 26 generation tasks based on the Big-Gen benchmark, demonstrating that our method requires only$\mathrm{O}(n)$human annotations for n iterations, whereas the score-based method requires$\mathrm{O}(n^{2})$annotations for the same number of iterations. Furthermore, we conducted at-test analysis, which indicated no statistically significant difference in effectiveness between our approach and the score-based method. These results suggest that our approach achieves comparable effectiveness while substantially reducing the annotation workload.

Read the paper · More papers on PaperTik