Towards Ordinal Data in LLM Evaluation Meta-analysis: A Non-parametric Perspective

Farouq Anbar, Andrei Mikriukov, Yaroslav Plaksin, Vladimir Sitnikov, Giancarlo Succi, Alexander G. Tormasov, Е. В. Трофимова · 2025

Ordinal data—data that are ordered categories but do not assume equal spacing between values—are firmly entrenched in social and biomedical research. Ordinal data, regardless of their prevalence, are most frequently analyzed through inappropriate parametric tests, leading to false results. The present study investigates appropriate statistical analysis of ordinal data on the basis of large language model (LLM) testing. From a collection of highly sophisticated LLM-created customized supplement regimens tested on ordinal scales by three GPT models, we report results using non-parametric testing in the form of the Mann-Whitney U test. Our findings establish statistically significant differences in evaluation models for many users, which provide evidence of variability in model judgment and necessity of proper ordinal data analysis. This article highlights the need for solid methodological precaution when performing ordinal scale analyses during AI testing.

Read the paper · More papers on PaperTik