🛡️
Session Flagged

Your session has been flagged for unusual activity.

You can try our app by searching for MultipleChat AI on Google and clicking the multiplechat.ai link to try it free.
Quick verification

Please confirm you're human to continue.


Frontier models, benchmarked continuously.

Quality, cost and speed for the leading AI models — refreshed on a live cycle · last updated 2026-09-13.

Intelligence, coding & agentic indexes

Composite indexes by Artificial Analysis · price per 1M tokens (input / output)

# Model Intelligence Coding Agentic $ / 1M in $ / 1M out
1 Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) 53.4 81.6 58 $5.00 $25.00
2 Qwen3.8 Max 53.4 68.9 49.9
3 GPT-6 Astra (max) 52.8 76.9 51.5 $11.00 $55.00
4 Claude Opus 5 (Adaptive Reasoning, Max Effort) 50.7 78 56.2 $5.50 $27.50
5 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) 49.7 76.5 51 $10.00 $50.00
6 GPT-5.6 Sol (max) 47.1 77.4 50.5 $4.00 $20.00
7 GLM-5.3 (max) 44.9 74.8 53.4 $1.40 $4.40
8 Grok 4.6 (high) 44.4 76.8 53.4 $4.00 $12.00
9 Kimi K3 (max) 43.8 76.2 50.6 $2.10 $10.95
10 GPT-5.6 Terra (max) 42.3 76.7 43.7 $4.00 $24.00
11 Claude Opus 4.8 (Adaptive Reasoning, Max Effort) 42 74.3 42.6 $10.00 $50.00
12 GLM-5.3-Flash 41.9 71.5 51.2 $0.070 $0.250
13 Gemini 3.8 Flash (high) 41.2 76.3 41.1 $1.35 $6.75
14 Qwen3.8 Max 40.3 71.8 49.6 $2.00 $6.00
15 Qwen3.8 2.4T A95B 40 71.9 50.4 $2.00 $6.00
16 Muse Spark 1.2 (xhigh) 39.8 72.2 44 $1.25 $4.25
17 DeepSeek V4.1 Flash (Reasoning, Max Effort) 39.5 $0.300 $1.20
18 Gemini 3.7 Flash (high) 39.4 76.1 36.4 $1.35 $6.75
19 Grok 4.5 (high) 39.1 72.4 42.1 $4.00 $12.00
20 GPT-5.5 (xhigh) 38.6 74.9 37.3 $2.50 $15.00
21 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) 38.4 71.5 44.3 $2.20 $11.00
22 GPT-5.6 Luna (max) 37.5 71.4 42.7 $0.400 $2.40
23 DeepSeek V4 Pro 0813 (Reasoning, Max Effort) 36.3 68.8 42.3 $0.880 $2.64
24 DeepSeek V4 Flash 0731 (Reasoning, Max Effort) 34.5 69.1 41.7 $0.110 $0.330
25 Muse Spark 1.1 (xhigh) 34.3 71.3 27.5 $1.25 $4.25
26 Gemini 3.6 Flash (high) 34.3 69.2 30.2 $1.35 $6.75
27 Qwen3.8 27B (xhigh) 33.9 68.1 46.5 $0.200 $2.50
28 Gemini 3.5 Flash (high) 33 70.1 27.3 $2.70 $16.20
29 DeepSeek V4 Pro (Reasoning, Max Effort) 30.9 59.4 27.7 $1.15 $2.55
30 Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) 30.5 63 33.1 $3.00 $15.00

Reasoning, agents & search — re-run on live models

GPQA Diamond

Graduate-level science reasoning

# Model Score $ / task
1 Gemini 3.1 Pro Preview 94.4% $0.202
2 GPT-6 Astra 94.4% $0.099
3 Gemini 3.7 Flash 94.3% $0.031
4 GPT-5.5 93.8% $0.308
5 GPT-5.6 Sol Pro 93.8% $0.445
6 Grok 4.6 93.3% $0.122
7 Gemini 3.6 Flash 92.8% $0.067
8 Gemini 3.5 Flash 92.8% $0.142
9 GPT-5.6 Sol 91.9% $0.071
10 Kimi K3 91.5% $0.147
11 MiniMax M3 90.6% $0.030
12 GPT-5.4 90.3% $0.147
13 GPT-5.6 Luna Pro 90.3% $0.044
14 DeepSeek V4.1 Flash 90.2% $0.032
15 Claude Fable 5.1 90.0% $0.127

No single model wins every benchmark.

The leader changes by task: one model tops reasoning, another tops coding, a third wins agentic search. That is exactly why MultipleChat exists — put the leaders side by side, let them draft, challenge and verify each other, and keep the best answer.

Quality, cost and speed together

A benchmark score without a price tag is marketing. Every leaderboard here pairs quality with cost per task or per million tokens.

Live & open

Refreshed continuously and downloadable as JSON. Our own quarterly comparison study lives at the MultipleChat AI Benchmark.

Attribution

Source: OpenRouter (openrouter.ai/rankings) incl. Artificial Analysis indexes, as of 2026-09-13T19:57:55Z. Licensed under CC BY 4.0.