Performance Analysis of Prompt-Engineering Techniques for Large Language Model

Miseol Son, Sungjin Lee · 2025

In this paper, we analyze the performance of three large language models (LLMs) - Llama3, Gemma2, and Mistral - using various combinations of prompt engineering techniques, focusing on the TruthfulQA and Winogrande datasets. In TruthfulQA, which emphasizes evaluating the truthfulness and cognitive errors of language models, the Chain of Thought (CoT) and Retrieval-Augmented Generation (RAG) techniques demonstrated superior performance. On the other hand, in Wino-grande, which assesses contextual reasoning abilities based on common-sense knowledge, CoT and In-Context Learning (ICL) were found to be effective. The performance analysis by model revealed that Llama3 exhibited the most significant improvement when prompt engineering techniques were applied. Meanwhile, Gemma2 achieved the highest overall performance across all evaluation metrics. This study highlights the importance of selecting appropriate prompt engineering techniques tailored to each model to maximize LLM performance, demonstrating that effective combinations of techniques can substantially enhance the models' reasoning abilities and accuracy.

Read the paper · More papers on PaperTik