Ranked by GPQA Diamond, graduate-level physics, chemistry and biology questions that require multi-step reasoning. This is the benchmark to watch when your product lives or dies on correctness.
Data updated: August 11, 2026
| # | Model | Vendor | Arena Elo | SWE-bench | Price in/out ($/M) | Context |
|---|---|---|---|---|---|---|
| 1 | Claude Mythos | Anthropic | 1478 | 93.9% | $25 / $125 | 1M |
| 2 | Claude Fable 5 | Anthropic | 1492 | 77% | $10 / $50 | 1M |
| 3 | Claude Mythos 5 | Anthropic | 1493 | 78% | $10 / $50 | 1M |
| 4 | Kimi K3 | Moonshot AI | 1478 | 75.5% | $3 / $15 | 1M |
| 5 | OpenAI o4 | OpenAI | 1438 | 72% | $12 / $48 | 256K |
| 6 | Gemini 3 Deep Think | Google DeepMind | 1429 | 64% | $14 / $56 | 1M |
| 7 | OpenAI o3 | OpenAI | 1418 | 69.1% | $2 / $8 | 200K |
| 8 | OpenAI o1 | OpenAI | 1380 | 48.9% | $15 / $60 | 200K |
| 9 | Grok 4 Heavy | xAI | 1391 | 55% | $15 / $60 | 256K |
| 10 | OpenAI o4-mini | OpenAI | 1362 | 60% | $1.1 / $4.4 | 200K |
| 11 | DeepSeek V4 | DeepSeek | 1395 | 62% | $1.74 / $3.48 | 1M |
| 12 | Gemini 3 Ultra | Google DeepMind | 1441 | 66% | $18 / $72 | 3M |
| 13 | DeepSeek R1 | DeepSeek | 1389 | 49.2% | $0.7 / $2.5 | 128K |
| 14 | GPT-5.5 | OpenAI | 1432 | 66% | $5 / $30 | 600K |
| 15 | Claude Opus 4.8 | Anthropic | 1435 | 67% | $5 / $25 | 1M |
GPQA Diamond is a set of PhD-level science questions written so that even experts with internet access find them hard. It is the cleanest signal we have for deep reasoning rather than memorization.
For multi-step problems where a wrong intermediate step ruins the answer, math, complex analysis, legal or scientific review, agent planning. For summaries, drafting and chat, a cheaper general model is usually enough.