Reference / 02
Who's actually
smarter.
Arena snapshot 2026-08-13 · 7.78M votes · 392 models
Arena Elo from 7.78M blind human votes, side by side with the benchmarks vendors love to quote — SWE-bench, GPQA, AIME, LiveCodeBench, MMMU, HLE. Every number links to its source; self-reported scores are flagged.
Models rated
10
▲ ranked by arena elo
Top arena elo
1456
▲ Claude Sonnet 4.5
Benchmarks tracked
20
▲ every score sourced
Snapshot
Aug 13, 2026
▲ 7.78M votes
The leaderboard
sorted by arena elo · click a column to re-sort— not measured / not published · † score has a caveat — hover the cell for details, see footnotes below.
Caveats †
- Qwen3 Max: all benchmark scores are vendor self-reported (Qwen), collected via the llm-stats aggregator — not independent measurements.
- Grok 4 SWE-bench Verified: third-party aggregators only, medium reliability — xAI did not publish an official SWE-bench score.
- GPT-5 mini LiveCodeBench v6 and MMLU-Pro: third-party measurement by LG AI (reasoning: high), not official OpenAI numbers.
Arena variants
- Claude Sonnet 4.5 rated as claude-sonnet-4-5-20250929-high-32k. base variant: 1455 Elo, rank 58.
- Claude Opus 4.1 rated as claude-opus-4-1-20250805-thinking-16k. base variant: 1447 Elo, rank 68.
- GPT-5 rated as gpt-5-high (reasoning). gpt-5-chat variant: 1427 Elo, rank 98.
- Qwen3 Max rated as qwen3-max-preview. only the preview variant is listed on the arena; final qwen3-max not found on the board.
- DeepSeek V3.2 rated as deepseek-v3.2 (Thinking).
- Gemini 2.5 Flash. preview-09-2025: 1404 Elo, rank 134.
- Grok 4 rated as grok-4-0709.
- GPT-5 mini rated as gpt-5-mini-high.
- Llama 4 Maverick rated as llama-4-maverick-17b-128e-instruct. April 2025 listing showed 1417 Elo for an experimental chat version that did not match the released weights ('Leaderboard Illusion'). Current rating reflects the public weights..
The Llama 4 «Leaderboard Illusion»
April 2025 listing showed 1417 Elo for an experimental chat version that did not match the released weights ('Leaderboard Illusion'). Current rating reflects the public weights.
Sources & methodology
Arena ratings
LMArena Text Overall leaderboard snapshot: 7.78M blind pairwise votes across 392 models, snapshot Aug 13, 2026. Arena Elo measures human preference in chat, not benchmark accuracy — a model can top one board and lag on the other. arena leaderboard ↗
Every number has a source
Benchmark scores are official vendor publications unless flagged †. Per-model source links:
- Claude Opus 4.1: anthropic.com ↗
- Claude Sonnet 4.5: anthropic.com ↗
- DeepSeek V3.2: arxiv.org ↗
- Gemini 2.5 Flash: arxiv.org ↗
- Gemini 2.5 Pro: arxiv.org ↗, deepmind.google ↗
- GPT-5 mini: arxiv.org ↗, github.com ↗
- GPT-5: openai.com ↗
- Grok 4: x.ai ↗, anotherwrapper.com ↗
- Llama 4 Maverick: huggingface.co ↗
- Qwen3 Max: llm-stats.com ↗