Frontier models, benchmarked continuously.
Quality, cost and speed for the leading AI models — refreshed on a live cycle · last updated 2026-10-08.
Intelligence, coding & agentic indexes
Composite indexes by Artificial Analysis · price per 1M tokens (input / output)
| # | Model | Intelligence | Coding | Agentic | $ / 1M in | $ / 1M out |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 (Max, Default Fallback) | 57.6 | — | — | $4.40 | $22.00 |
| 2 | Claude Sonnet 5.5 (Max, Default Fallback) | 56 | — | — | $2.00 | $10.00 |
| 3 | Claude Fable 5.1 (Max, Default Fallback) | 53.4 | 81.6 | 57.9 | $5.00 | $25.00 |
| 4 | Qwen3.8 Max | 53.4 | 68.9 | 49.9 | — | — |
| 5 | GPT-6 Astra (Max) | 52.7 | 76.9 | 51 | $60.00 | $300.00 |
| 6 | GPT-6.1 Sol (Max) | 51.8 | — | — | $2.20 | $11.00 |
| 7 | Claude Opus 5 (Max) | 50.8 | 78 | 56.5 | $5.50 | $27.50 |
| 8 | Claude Fable 5 (Max, Opus 4.8 Fallback) | 49.6 | 76.5 | 50.7 | $10.00 | $50.00 |
| 9 | Muse Spark 1.3 (Max) | 48.1 | 75.8 | 55.5 | $1.25 | $4.25 |
| 10 | GPT-6 Sol (Max) | 47.6 | — | — | $2.00 | $10.00 |
| 11 | GPT-5.6 Sol (Max) | 47 | 77.4 | 50.2 | $4.00 | $20.00 |
| 12 | Grok 4.7 (Xhigh) | 46.4 | — | — | $2.20 | $6.60 |
| 13 | MiMo-V2.6-Pro | 46.3 | — | — | $0.430 | $0.870 |
| 14 | Qwen3.8 Max (0902) | 45.4 | 76.2 | 56 | $2.00 | $6.00 |
| 15 | GLM-5.3 (Max) | 44.8 | 74.8 | 53.1 | $2.10 | $6.60 |
| 16 | Grok 4.6 (High) | 44.3 | 76.8 | 53 | $2.20 | $6.60 |
| 17 | Kimi K3 (Max) | 43.6 | 76.2 | 50 | $2.99 | $15.00 |
| 18 | Claude Haiku 5.5 (Max) | 43.4 | — | — | $0.110 | $0.550 |
| 19 | GPT-5.6 Terra (Max) | 42.1 | 76.7 | 43.2 | $4.00 | $24.00 |
| 20 | Claude Opus 4.8 (Max) | 41.8 | 74.3 | 41.9 | $10.00 | $50.00 |
| 21 | GLM 5.3 Flash | 41.8 | 71.5 | 50.9 | $0.190 | $0.620 |
| 22 | Ling 3.1 Flash | 41.1 | — | — | $0 | $0 |
| 23 | Gemini 3.8 Flash (High) | 40.9 | 76.3 | 40.2 | $1.35 | $6.75 |
| 24 | Claude Opus 4.7 (Max) | 40.7 | 73.6 | 38.6 | $5.00 | $25.00 |
| 25 | Qwen3.8 2.4T A95B | 39.9 | 71.9 | 50.1 | $2.00 | $6.00 |
| 26 | Muse Spark 1.2 (Xhigh) | 39.6 | 72.2 | 43.2 | $1.25 | $4.25 |
| 27 | Gemini 3.7 Flash (Medium) | 39.6 | 71.5 | — | $1.35 | $6.75 |
| 28 | DeepSeek V4.1 Flash (Max) | 39.5 | — | — | $0.090 | $0.180 |
| 29 | GPT-5.4 (Xhigh) | 39 | 71.1 | — | $5.00 | $30.00 |
| 30 | Grok 4.5 (High) | 38.8 | 72.4 | 41.2 | $4.00 | $12.00 |
Reasoning, agents & search — re-run on live models
GPQA Diamond
Graduate-level science reasoning
| # | Model | Score | $ / task |
|---|---|---|---|
| 1 | Gemini 3.8 Flash | 95.6% | $0.068 |
| 2 | GPT-6 Astra Pro | 95.5% | $0.333 |
| 3 | Fugu Ultra | 94.6% | $1.15 |
| 4 | Gemini 3.1 Pro Preview | 94.5% | $0.202 |
| 5 | GPT-6 Astra | 94.4% | $0.113 |
| 6 | GPT-6.1 Sol | 94.4% | $0.021 |
| 7 | Jev Router | 93.9% | $0.016 |
| 8 | GPT-5.6 Sol Pro | 93.9% | $0.298 |
| 9 | Gemini 3.7 Flash | 93.9% | $0.029 |
| 10 | GPT-5.5 | 93.7% | $0.309 |
| 11 | Fugu Ultra V2 | 93.3% | $0.503 |
| 12 | Grok 4.6 | 92.8% | $0.123 |
| 13 | Gemini 3.6 Flash | 92.6% | $0.060 |
| 14 | Gemini 3.5 Flash | 92.6% | $0.141 |
| 15 | Pareto | 92.4% | $0.034 |
No single model wins every benchmark.
The leader changes by task: one model tops reasoning, another tops coding, a third wins agentic search. That is exactly why MultipleChat exists — put the leaders side by side, let them draft, challenge and verify each other, and keep the best answer.
Quality, cost and speed together
A benchmark score without a price tag is marketing. Every leaderboard here pairs quality with cost per task or per million tokens.
Live & open
Refreshed continuously and downloadable as JSON. Our own quarterly comparison study lives at the MultipleChat AI Benchmark.
Attribution
Source: OpenRouter (openrouter.ai/rankings) incl. Artificial Analysis indexes, as of 2026-10-08T08:19:28Z. Licensed under CC BY 4.0.