Real-world Evaluation of Modern Generative AI Systems

Angeliki Romanou · Infoscience (Ecole Polytechnique Fédérale de Lausanne) · 2026

Modern generative AI systems are expected to operate across diverse tasks, languages, users, and deployment settings, reflecting the heterogeneity and variability of real-world use. Yet current evaluation practices often remain disconnected from the conditions under which these systems are actually used. Standard AI evaluation benchmarks typically rely on controlled inputs, narrow linguistic and regional coverage, and fixed evaluation protocols, while real-world use involves diverse user communities, region-specific knowledge, noisy and variable inputs, and context-dependent interactions. Designing evaluations that better reflect these conditions is only part of the challenge, as such evaluations must also be incorporated throughout model development so that decisions made during training and adaptation are guided by evidence that remains relevant to eventual deployment. This thesis studies how evaluation can better support claims about the capabilities and limitations of modern AI systems under real-world conditions. It develops benchmarks, analysis frameworks, and evaluation methodologies around three complementary requirements. Evaluations should be inclusive of diverse linguistic and regional environments, robust to variability and uncertainty in user interactions, and phase-aware so that they support model assessment throughout development and toward deployment. First, we broaden multilingual evaluation beyond translated benchmarks by developing large-scale evaluations grounded in native academic, professional, and occupational examinations across 91 languages and 86 countries, together with a framework that distinguishes linguistic ability from access to regional knowledge; our results reveal substantial disparities across languages, regions, and knowledge types, particularly for locally specific knowledge. Second, we study the robustness and uncertainty of evaluation conclusions by measuring sensitivity to semantics-preserving prompt variations and evaluating graded, context-dependent causal reasoning over real-world events, showing that both benchmark measurements and model behavior can vary substantially under conditions not captured by standard evaluation protocols. Third, we develop phase-aware, end-to-end evaluation methodologies that integrate evaluation throughout model development and toward deployment, using evaluation to monitor capability acquisition, inform development decisions, assess interventions, and characterize final models across capabilities, languages, and regional contexts. We extend this approach to specialized domains by combining standardized benchmarks with expert-driven, open-ended evaluation of deployment readiness, demonstrating the importance of adapting evaluation methods to both the stage of model development and the conditions of intended real-world use. The main contributions of this thesis include large-scale multilingual and regional knowledge evaluation, methodologies for measuring benchmark robustness and uncertainty, and phase-aware evaluation frameworks that integrate evaluation throughout model development and deployment-readiness assessment. Together, these contributions provide practical foundations for evaluating generative AI systems across diverse real-world contexts, interaction variability, and stages of development.

Read the paper · More papers on PaperTik