Ranked by Chatbot Arena Elo. The community-driven score where models face off blind head-to-head. Higher Elo means humans consistently prefer its answers in direct comparison.
Data updated: August 11, 2026
| # | Model | Vendor | Arena Elo | SWE-bench | Price in/out ($/M) | Context |
|---|---|---|---|---|---|---|
| 1 | Claude Mythos 5 | Anthropic | 1493 | 78% | $10 / $50 | 1M |
| 2 | Claude Fable 5 | Anthropic | 1492 | 77% | $10 / $50 | 1M |
| 3 | Claude Mythos | Anthropic | 1478 | 93.9% | $25 / $125 | 1M |
| 4 | Kimi K3 | Moonshot AI | 1478 | 75.5% | $3 / $15 | 1M |
| 5 | Gemini 3 Ultra | Google DeepMind | 1441 | 66% | $18 / $72 | 3M |
| 6 | OpenAI o4 | OpenAI | 1438 | 72% | $12 / $48 | 256K |
| 7 | Claude Opus 4.8 | Anthropic | 1435 | 67% | $5 / $25 | 1M |
| 8 | GPT-5.5 | OpenAI | 1432 | 66% | $5 / $30 | 600K |
| 9 | Gemini 3 Deep Think | Google DeepMind | 1429 | 64% | $14 / $56 | 1M |
| 10 | Claude Opus 4.7 | Anthropic | 1420 | 64.3% | $5 / $25 | 1M |
| 11 | Claude Code | Anthropic | 1420 | 64.3% | $15 / $75 | 1M |
| 12 | OpenAI o3 | OpenAI | 1418 | 69.1% | $2 / $8 | 200K |
| 13 | GPT-5 | OpenAI | 1412 | 65% | $1.25 / $10 | 400K |
| 14 | GPT-5.4 Codex | OpenAI | 1408 | 70% | $9 / $36 | 500K |
| 15 | Codex CLI | OpenAI | 1408 | 70% | $9 / $36 | 500K |
By community preference (Arena Elo) the top is held by frontier models from Anthropic, OpenAI and Google. The practical answer depends on your use case, coding, reasoning and price each produce a different winner, which is why we publish one leaderboard per dimension.
Chatbot Arena shows people two anonymous answers to the same prompt and asks which is better. Each vote updates an Elo rating, the same system used in chess. It is the closest thing to a blind taste test for AI models.