Independent, weekly-updated rankings and head-to-head comparisons of the top AI models (LLMs). Built by engineers who ship production AI agents.
Data updated: August 11, 2026
| # | Model | Vendor | Arena Elo | SWE-bench | Price in/out ($/M) | Context |
|---|---|---|---|---|---|---|
| 1 | Claude Mythos 5 | Anthropic | 1493 | 78% | $10 / $50 | 1M |
| 2 | Claude Fable 5 | Anthropic | 1492 | 77% | $10 / $50 | 1M |
| 3 | Claude Mythos | Anthropic | 1478 | 93.9% | $25 / $125 | 1M |
| 4 | Kimi K3 | Moonshot AI | 1478 | 75.5% | $3 / $15 | 1M |
| 5 | Gemini 3 Ultra | Google DeepMind | 1441 | 66% | $18 / $72 | 3M |
| 6 | OpenAI o4 | OpenAI | 1438 | 72% | $12 / $48 | 256K |
| 7 | Claude Opus 4.8 | Anthropic | 1435 | 67% | $5 / $25 | 1M |
| 8 | GPT-5.5 | OpenAI | 1432 | 66% | $5 / $30 | 600K |
| 9 | Gemini 3 Deep Think | Google DeepMind | 1429 | 64% | $14 / $56 | 1M |
| 10 | Claude Opus 4.7 | Anthropic | 1420 | 64.3% | $5 / $25 | 1M |
| 11 | Claude Code | Anthropic | 1420 | 64.3% | $15 / $75 | 1M |
| 12 | OpenAI o3 | OpenAI | 1418 | 69.1% | $2 / $8 | 200K |
| 13 | GPT-5 | OpenAI | 1412 | 65% | $1.25 / $10 | 400K |
| 14 | GPT-5.4 Codex | OpenAI | 1408 | 70% | $9 / $36 | 500K |
| 15 | Codex CLI | OpenAI | 1408 | 70% | $9 / $36 | 500K |
Estimate your monthly API bill with the LLM cost calculator
Public benchmarks are almost always measured in English, and that hides real differences, frontier models (Claude, GPT, Gemini) keep near-identical quality in Spanish, while small and open-weight models tend to lose more precision on nuance, regional idioms and formal writing.
At Alher Tech we evaluate models on real Spanish-language tasks (customer support, legal documents, marketing copy) before recommending them to a client. If your product speaks Spanish, always test with your own Spanish prompts, the English ranking does not always hold.
There is no single winner, it depends on the workload. By Chatbot Arena Elo the top of the ranking is led by frontier models like Claude Fable 5, GPT-5.5 and Gemini 3 Ultra, but for coding, reasoning or budget workloads the ranking changes. Our leaderboards sort every model by the metric that matters for each use case.
For code generation we rank models by HumanEval and SWE-bench Verified, the benchmarks that measure real bug-fixing on production repositories. Check the "Best AI for Coding" leaderboard for the current top, claude and GPT frontier models usually lead, with DeepSeek as the strongest open-weights option.
Every model is scored with the same public benchmarks (Chatbot Arena Elo, MMLU, GPQA Diamond, HumanEval, SWE-bench, AIME) plus its official list pricing and context window. Data comes from the benchmark publishers and each vendor’s documentation, and each model card shows its own last-updated date.
The "Cheapest LLMs" leaderboard ranks models by blended price per million tokens. Small models like Claude Haiku, GPT-5 mini or Gemini Flash cost a fraction of the flagships and are often enough for classification, extraction and support chatbots.
Benchmark and pricing data is refreshed on every deploy from public sources, and each model card carries its own last-updated date so you can see exactly how fresh its numbers are.