Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Tu Vu, Kalpesh Krishna, Salaheddin Alzubi, Chris Tar, Manaal Faruqui, Yun-Hsuan Sung · 2024

As large language models (LLMs) evolve, evaluating their output reliably becomes increasingly difficult due to the high cost of human evaluation.To address this, we introduce FLAMe, a family of Foundational Large Autorater Models.FLAMe is trained on a diverse set of over 100 quality assessment tasks, incorporating 5M+ human judgments curated from publicly released human evaluations.FLAMe outperforms models like GPT-4 and Claude-3 on various held-out tasks, and serves as a powerful starting point for finetuning, as shown in our reward model evaluation case study (FLAMe-RM).On Reward-Bench, FLAMe-RM-24B achieves 87.8% accuracy, surpassing GPT-4-0125 (85.9%) and GPT-4o (84.7%).Additionally, we introduce FLAMe-Opt-RM, an efficient tail-patch finetuning approach that offers competitive Re-wardBench performance using 25× fewer training datapoints.Our FLAMe variants outperform popular proprietary LLM-as-a-Judge models on 8 of 12 autorater benchmarks, covering 53 quality assessment tasks, including RewardBench and LLM-AggreFact.Finally, our analysis shows that FLAMe is significantly less biased than other LLM-as-a-Judge models on the CoBBLEr autorater bias benchmark.1 * Tu Vu and Kalpesh Krishna contributed equally to the project leadership, design, and implementation of the work.† Work done while at UMass Amherst.‡ Equal contribution as senior advisors.1 The FLAMe collection is available at https:// huggingface.co/datasets/google/flame-collection."""Input format.""" INSTRUCTIONS:"""Task definition and evaluation instructions."""

Read the paper · More papers on PaperTik