Kimi K3
Last updated: July 18, 2026
Specs
| Vendor | Moonshot AI |
|---|
| Released | July 17, 2026 |
|---|
| Context window | 1M tokens |
|---|
| Input | $3 / M tokens |
|---|
| Output | $15 / M tokens |
|---|
Benchmarks
| mmlu | 94 |
|---|
| gpqa | 87.5 |
|---|
| humanEval | 96.5 |
|---|
| arenaElo | 1478 |
|---|
| sweBench | 75.5 |
|---|
| sweBenchPro | 70 |
|---|
| liveCodeBench | 84.5 |
|---|
| aiderPolyglot | 83 |
|---|
| terminalBench | 49 |
|---|
| mmluPro | 88 |
|---|
| hle | 26 |
|---|
| aime | 92 |
|---|
| math500 | 97 |
|---|
| mmmu | 80 |
|---|
Strengths
- Largest open-weight model ever (2.8T MoE, 16 of 896 experts). Frontier-class results at a third of closed-flagship pricing
- 1M context with native text, image and video understanding and always-on reasoning
- Weights due July 27, 2026 under a modified MIT license. Self-hosting and EU data residency become possible
Weaknesses
- Most launch benchmarks are self-reported; independent replication has barely started
- Moonshot acknowledges excessive proactiveness in task execution and sensitivity to preserved thinking history
- Serving a 2.8T MoE yourself demands serious infrastructure; hosted capacity outside China is still ramping
Best for
- Cost-sensitive agentic coding and long-context workloads near the frontier
- Teams that need open weights for compliance, data residency or on-prem deployment
- High-volume products routing bulk traffic away from premium closed flagships
Compare with other AI models