Official AI Models & LLM Benchmark Hub (2026 Live Verified)

AI Models & LLM Intelligence Leaderboard

Comprehensive, zero-drift AI benchmarks aggregated across LMSYS Arena.ai (human Elo), Artificial Analysis (quality/speed ratio & price), SWE-bench Verified, GPQA Diamond, LiveBench, and MTEB.

#1 Frontier Reasoning
Claude Opus 5 (High)
Anthropic • Ultra-MoE
1504 Arena Elo • 78.6% SWE
#1 Test-Time Compute
GPT-5.6 Sol (xHigh)
OpenAI • Adaptive MoE
1498 Arena Elo • 97.6 Intel
#1 Quality / Speed Ratio
Gemini 3.7 Flash
Google • 2M Context
158 pts Ratio • 165 tps ($0.61/M)
#1 Open Weights MIT
DeepSeek-V4 Pro
DeepSeek • 750B MoE
1474 Arena Elo • $0.35/M
License: All Open Proprietary
# Model & Architecture Arena Elo (Human) SWE / GPQA Speed & Ratio Price ($/1M) Context Actions
1
Claude Opus 5 (High)
Anthropic Ultra-MoE proprietary
1504 (27.6k)
Agent: 1535
78.6% SWE
GPQA: 84.2%
84 tps 83 pts
TTFT: 460ms
$10.00
Cache: $0.50
500k
Tokens
2
GPT-5.6 Sol (xHigh)
OpenAI Omni-MoE proprietary
1498 (42.1k)
Agent: 1520
76.8% SWE
GPQA: 83.1%
96 tps 94 pts
TTFT: 380ms
$9.00
Cache: $0.45
400k
Tokens
3
Gemini 3.7 Flash (High)
Google DeepMind Fast-MoE proprietary
1490 (5.7k)
Agent: 1492
73.5% SWE
GPQA: 80.4%
165 tps 158 pts
TTFT: 190ms
$0.61
Cache: $0.08
2.0M
Tokens
4
Muse-Spark 1.2 (xHigh)
Meta Spark-MoE proprietary
1487 (3.3k)
72.1% SWE
GPQA: 79.5%
125 tps 119 pts
TTFT: 240ms
$1.75
Blended 3:1
500k
Tokens
5
Grok 4.6 (Thinking)
xAI Grok-MoE proprietary
1480 (14.2k)
69.5% SWE
GPQA: 77.1%
92 tps 86 pts
TTFT: 340ms
$4.00
Blended 3:1
256k
Tokens
6
DeepSeek-V4 Pro (High)
DeepSeek 750B (40B active) 64GB open-source
1474 (16.8k)
Agent: 1480
74.2% SWE
GPQA: 80.8%
78 tps 73 pts
TTFT: 410ms
$0.61
Blended 3:1
256k
Tokens
7
DeepSeek-R1 (RL)
DeepSeek 671B (37B active) 48GB open-source
1458 (18.5k)
70.2% SWE
GPQA: 78.4%
62 tps 57 pts
TTFT: 650ms
$0.96
Blended 3:1
128k
Tokens

Query this Leaderboard via REST API & Agents

Instant, zero-auth JSON endpoint with edge caching for programmatic model selection.

GET https://yakaai.com/api/leaderboard