Study of 160 commonsense problems found that reasoning models and humans fail on the same questions, and that averaging chain-of-thought length across multiple runs raises the human-model difficulty correlation from 0.41 to 0.55 for GPT-OSS-20B
LLM benchmark rankings and evaluation setup
LLM Rankings Shift With Evaluation Setup, Data Quality, and Run Count
Benchmark rankings widely used to compare models are sensitive to evaluation setup choices, data quality, and how many runs are averaged, not just to differences in model capability. Rohan paul concluded that no single evaluation setup should decide a leaderboard, and a commonsense reasoning study found that single-run chain-of-thought length is a noisy difficulty signal that improves substantially with multi-run averaging.
The full picture
Three lines of research show that LLM benchmark rankings reflect evaluation methodology as much as model capability. Rohan paul found that varying only evaluation parameters, such as prompt format, option order, and scoring method, across 3,679 questions and 12 fixed models caused Gemma 4-31B's score to range from 31% to 89%, with four of twelve models reaching rank 1 under at least one valid configuration. The scoring method was the single biggest driver, with 95.7% of the average gap between neighboring models coming from questions whose answers change when the evaluation setup changes. Cameron Wolfe's overview of benchmark methods notes a parallel effect from data quality: removing incorrect samples from MMLU caused Llama-3.1-405B to rise from 16th to 1st in Virology and Qwen-2-72B-Instruct to drop from 1st to 8th in College Chemistry. As a partial remedy, the overview describes Fluid Benchmarking, which dynamically selects the most informative evaluation examples for a particular model. A study of 160 commonsense problems found that reasoning models and humans tend to fail on the same questions, and that chain-of-thought length as a difficulty signal becomes more reliable when averaged across multiple runs; for GPT-OSS-20B, the human-model difficulty correlation rose from 0.41 to 0.55 when averaged across multiple reasoning paths.
How it developed
Two studies published September 2 found that scaffold and prompt choices used to run LLM benchmarks contribute nearly as much variance to scores as model differences.
Zhang et al. on SWE-bench Verified found harness-induced variance 7.8x model-induced; a second study of 3,679 questions found 95.7% of the gap between neighboring models came from setup-sensitive questions. On September 3, analyst Vaughan's survey of three benchmarks put the harness effect at 27.4 points of spread against a 29.4-point model effect.
FM-Bench shows year-5 rankings correlate just 0.19 with final order in 20-year simulated management task; token usage varied 7x but did not predict ranking
Additional coverage of Zhang et al. harness paper, specifying pass@1 ranges: 8.5-13.0 pp for harness switch vs. 2.5-5.0 pp for model switch
Study found that varying only evaluation parameters across 3,679 questions and 12 fixed models caused large ranking swings, with Gemma 4-31B ranging from 31% to 89% accuracy
WebDev-Skills-Bench study reports skill file injection lowers pass rates and raises token costs across all tested models
Sources
- Reasoning models struggle on many of the same problems humans do,
- Current agent benchmarks may be ending before the real failures start.
- OpenClaw 2.0 vs Hermes experiment by @atomicbot_ai is a good example of why the model alone tells you very little about …
- For long-horizon agents, this paper argues the harness can matter more than the model, so benchmark scores should not be…
- LLM rankings can be created by evaluation choices as much as model differences, so one setup should never decide the lea…
- Given 6 days and $3K AI agents, produced 2 research papers, and both were rejected.
- Keep the model fixed, change the harness, and coding-agent results can move a lot when context gets tight.
- Attaching a skill file to every prompt in a coding session is usually a net loss.
1 more source
Want this in your inbox?
I send a short email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free