Consistency evaluation protocol: A reproducible framework for assessing large language model output repeatability

Shraddha Vaidya, Jatinderkumar R. Saini · MethodsX · 2026

The stochastic behaviour of Large Language Models (LLMs) generates varied responses when prompted with the same inputs and parameters. Recent researchers have explored various techniques to understand this behaviour of LLMs, including uncertainty measures, semantic consistency, robustness, and prompt sensitivity. However, the research lacks a reproducible methodology that combines semantic and structural parameters. Thus, the proposed consistency evaluation protocol focuses on evaluating the consistency of LLM responses through repeated prompting, extracting relevant features, and analysing structural and semantic stability through consistency scores. The framework was evaluated using Mistral-7B-Instruct-v0.2 on a set of 20 input texts. For each input text, five independent responses were generated. For each of the generated responses, semantic consistency is determined through embedding-based text representations. In addition, the structural variation is studied through fluctuations in sentence length, word usage, and lexical diversity. Lastly, these measurements are combined to form the Composite Consistency Score (CCS). Further, the validation tests are performed, containing temperature sensitivity, prompt sensitivity, and metric validations reflecting reliable and reproducible consistency evaluations.

Read the paper · More papers on PaperTik