🛡️
Session Flagged

Your session has been flagged for unusual activity.

You can try our app by searching for MultipleChat AI on Google and clicking the multiplechat.ai link to try it free.
Quick verification

Please confirm you're human to continue.


Frontier models, benchmarked continuously.

Quality, cost and speed for the leading AI models — refreshed on a live cycle · last updated 2026-10-08.

Intelligence, coding & agentic indexes

Composite indexes by Artificial Analysis · price per 1M tokens (input / output)

# Model Intelligence Coding Agentic $ / 1M in $ / 1M out
1 Claude Opus 5.5 (Max, Default Fallback) 57.6 — — $4.40 $22.00
2 Claude Sonnet 5.5 (Max, Default Fallback) 56 — — $2.00 $10.00
3 Claude Fable 5.1 (Max, Default Fallback) 53.4 81.6 57.9 $5.00 $25.00
4 Qwen3.8 Max 53.4 68.9 49.9 — —
5 GPT-6 Astra (Max) 52.7 76.9 51 $60.00 $300.00
6 GPT-6.1 Sol (Max) 51.8 — — $2.20 $11.00
7 Claude Opus 5 (Max) 50.8 78 56.5 $5.50 $27.50
8 Claude Fable 5 (Max, Opus 4.8 Fallback) 49.6 76.5 50.7 $10.00 $50.00
9 Muse Spark 1.3 (Max) 48.1 75.8 55.5 $1.25 $4.25
10 GPT-6 Sol (Max) 47.6 — — $2.00 $10.00
11 GPT-5.6 Sol (Max) 47 77.4 50.2 $4.00 $20.00
12 Grok 4.7 (Xhigh) 46.4 — — $2.20 $6.60
13 MiMo-V2.6-Pro 46.3 — — $0.430 $0.870
14 Qwen3.8 Max (0902) 45.4 76.2 56 $2.00 $6.00
15 GLM-5.3 (Max) 44.8 74.8 53.1 $2.10 $6.60
16 Grok 4.6 (High) 44.3 76.8 53 $2.20 $6.60
17 Kimi K3 (Max) 43.6 76.2 50 $2.99 $15.00
18 Claude Haiku 5.5 (Max) 43.4 — — $0.110 $0.550
19 GPT-5.6 Terra (Max) 42.1 76.7 43.2 $4.00 $24.00
20 Claude Opus 4.8 (Max) 41.8 74.3 41.9 $10.00 $50.00
21 GLM 5.3 Flash 41.8 71.5 50.9 $0.190 $0.620
22 Ling 3.1 Flash 41.1 — — $0 $0
23 Gemini 3.8 Flash (High) 40.9 76.3 40.2 $1.35 $6.75
24 Claude Opus 4.7 (Max) 40.7 73.6 38.6 $5.00 $25.00
25 Qwen3.8 2.4T A95B 39.9 71.9 50.1 $2.00 $6.00
26 Muse Spark 1.2 (Xhigh) 39.6 72.2 43.2 $1.25 $4.25
27 Gemini 3.7 Flash (Medium) 39.6 71.5 — $1.35 $6.75
28 DeepSeek V4.1 Flash (Max) 39.5 — — $0.090 $0.180
29 GPT-5.4 (Xhigh) 39 71.1 — $5.00 $30.00
30 Grok 4.5 (High) 38.8 72.4 41.2 $4.00 $12.00

Reasoning, agents & search — re-run on live models

GPQA Diamond

Graduate-level science reasoning

# Model Score $ / task
1 Gemini 3.8 Flash 95.6% $0.068
2 GPT-6 Astra Pro 95.5% $0.333
3 Fugu Ultra 94.6% $1.15
4 Gemini 3.1 Pro Preview 94.5% $0.202
5 GPT-6 Astra 94.4% $0.113
6 GPT-6.1 Sol 94.4% $0.021
7 Jev Router 93.9% $0.016
8 GPT-5.6 Sol Pro 93.9% $0.298
9 Gemini 3.7 Flash 93.9% $0.029
10 GPT-5.5 93.7% $0.309
11 Fugu Ultra V2 93.3% $0.503
12 Grok 4.6 92.8% $0.123
13 Gemini 3.6 Flash 92.6% $0.060
14 Gemini 3.5 Flash 92.6% $0.141
15 Pareto 92.4% $0.034

No single model wins every benchmark.

The leader changes by task: one model tops reasoning, another tops coding, a third wins agentic search. That is exactly why MultipleChat exists — put the leaders side by side, let them draft, challenge and verify each other, and keep the best answer.

Quality, cost and speed together

A benchmark score without a price tag is marketing. Every leaderboard here pairs quality with cost per task or per million tokens.

Live & open

Refreshed continuously and downloadable as JSON. Our own quarterly comparison study lives at the MultipleChat AI Benchmark.

Attribution

Source: OpenRouter (openrouter.ai/rankings) incl. Artificial Analysis indexes, as of 2026-10-08T08:19:28Z. Licensed under CC BY 4.0.