Comparative Study of Large Language Model Evaluation Frameworks with a Focus on NLP vs LLM-As-A-Judge Metrics
Afnan Alabdulwahab, Chloe Japic, Chao Le, Disha Dubey, Disha Trivedi, John Hope, Patrick Stone, Sanjana Srivastava, Adam P. Tashman, Aidong Zhang · 2025
As Large Language Models (LLMs) are increasingly deployed across industries, evaluating their performance remains a core challenge. In collaboration with Deloitte, this study compares traditional Natural Language Processing (NLP) metrics with the emerging LLM-as-a-Judge paradigm across tasks, including retrieval, response accuracy, toxicity, bias, hallucination, summarization, tone, and readability. Using models like Claude and both public and synthetic datasets, we conduct structured evaluations via API. Results show that LLM-as-a-Judge methods offer nuanced assessments but face self-referential bias and consistency issues, while traditional metrics remain transparent yet limited. We propose an evaluation playbook outlining tradeoffs and best-use scenarios to support standardized, responsible LLM assessment.