metabench - Paper Data

Alex Kipnis, Konstantinos Voudouris, Luca M. Schulze Buschoff, Eric R. Schulz · arXiv (Cornell University) · 2024

Item-wise accuracies in six benchmarks from Open LLM Leaderboard 1 scraped from huggingface.co and used for metabench analyses and construction. Datasets with RMSE's for random benchmark subsets are used as reference in the paper and are included here.

Read the paper · More papers on PaperTik