The Information Machine
Following·since 31 Aug 2026·Day 4·6 sources·updated 3 Sep 2026

LLM benchmark rankings and evaluation setup

The gist

LLM Rankings Shift With Evaluation Setup, Data Quality, and Run Count

Benchmark rankings widely used to compare models are sensitive to evaluation setup choices, data quality, and how many runs are averaged, not just to differences in model capability. Rohan paul concluded that no single evaluation setup should decide a leaderboard, and a commonsense reasoning study found that single-run chain-of-thought length is a noisy difficulty signal that improves substantially with multi-run averaging.

The full picture

Three lines of research show that LLM benchmark rankings reflect evaluation methodology as much as model capability. Rohan paul found that varying only evaluation parameters, such as prompt format, option order, and scoring method, across 3,679 questions and 12 fixed models caused Gemma 4-31B's score to range from 31% to 89%, with four of twelve models reaching rank 1 under at least one valid configuration. The scoring method was the single biggest driver, with 95.7% of the average gap between neighboring models coming from questions whose answers change when the evaluation setup changes. Cameron Wolfe's overview of benchmark methods notes a parallel effect from data quality: removing incorrect samples from MMLU caused Llama-3.1-405B to rise from 16th to 1st in Virology and Qwen-2-72B-Instruct to drop from 1st to 8th in College Chemistry. As a partial remedy, the overview describes Fluid Benchmarking, which dynamically selects the most informative evaluation examples for a particular model. A study of 160 commonsense problems found that reasoning models and humans tend to fail on the same questions, and that chain-of-thought length as a difficulty signal becomes more reliable when averaged across multiple runs; for GPT-OSS-20B, the human-model difficulty correlation rose from 0.41 to 0.55 when averaged across multiple reasoning paths.

How it developed
5 September 2026

Study of 160 commonsense problems found that reasoning models and humans fail on the same questions, and that averaging chain-of-thought length across multiple runs raises the human-model difficulty correlation from 0.41 to 0.55 for GPT-OSS-20B

3 September 2026

Two studies published September 2 found that scaffold and prompt choices used to run LLM benchmarks contribute nearly as much variance to scores as model differences.

Zhang et al. on SWE-bench Verified found harness-induced variance 7.8x model-induced; a second study of 3,679 questions found 95.7% of the gap between neighboring models came from setup-sensitive questions. On September 3, analyst Vaughan's survey of three benchmarks put the harness effect at 27.4 points of spread against a 29.4-point model effect.

2 September 2026

FM-Bench shows year-5 rankings correlate just 0.19 with final order in 20-year simulated management task; token usage varied 7x but did not predict ranking

1 September 2026

Additional coverage of Zhang et al. harness paper, specifying pass@1 ranges: 8.5-13.0 pp for harness switch vs. 2.5-5.0 pp for model switch

31 August 2026

Study found that varying only evaluation parameters across 3,679 questions and 12 fixed models caused large ranking swings, with Gemma 4-31B ranging from 31% to 89% accuracy

26 August 2026

WebDev-Skills-Bench study reports skill file injection lowers pass rates and raises token costs across all tested models

Sources
1 more source
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free