Together AI published 904-rollout DeepSWE comparison of Claude Fable 5 and DeepSeek V4 Pro 0813, finding Fable 5 leads pass@1 at roughly 90x the cost while DeepSeek wins pass@4; cascade routing hits 82.7%
Claude Fable 5 leads GPT-5.5 on math and DeepSeek on coding pass@1
FrontierMath is designed to be saturation-resistant, so a 13-point gap between models signals genuine capability differences rather than benchmark overfitting. On coding, the 90x cost difference means the choice between Fable 5 and DeepSeek V4 Pro turns on whether single-attempt accuracy or multi-attempt cost-efficiency matters more for a given use case.
The full picture
Claude Fable 5 scored 88% on FrontierMath tier 4 under Epoch AI evaluation, beating GPT-5.5's roughly 75% by 13 percentage points. For context, Anthropic's Opus 4.5 scored below 10% on the same tier earlier in 2026. On the DeepSWE coding benchmark, Together AI ran 904 rollouts and found Fable 5 leads DeepSeek V4 Pro 0813 on pass@1 accuracy but costs approximately 90 times more; DeepSeek V4 Pro wins on pass@4. A cascade routing strategy combining both models reaches 82.7% on DeepSWE.
How it developed
Sources
Want this in your inbox?
I send a short email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free