no-way.dev

Reference / 02

Who's actually
smarter.

Arena snapshot 2026-08-13 · 7.78M votes · 392 models

Arena Elo from 7.78M blind human votes, side by side with the benchmarks vendors love to quote — SWE-bench, GPQA, AIME, LiveCodeBench, MMMU, HLE. Every number links to its source; self-reported scores are flagged.

Models rated

10

▲ ranked by arena elo

Top arena elo

1456

▲ Claude Sonnet 4.5

Benchmarks tracked

20

▲ every score sourced

Snapshot

Aug 13, 2026

▲ 7.78M votes

The leaderboard

sorted by arena elo · click a column to re-sort
01Closedrank 55swe 77.2gpqa 83.4aime 87.0
02Closedrank 64swe 74.5gpqa 80.9aime 78.0
03Closedrank 70swe 63.2gpqa 86.4aime 88.0
GPT-51434
04Closedrank 87swe 74.9gpqa 88.4aime 94.6
05Closedrank 86swe 69.6gpqa 62.0aime 81.6
06Openrank 101swe 73.1gpqa 82.4aime 93.1
07Closedrank 129swe 48.9gpqa 82.8aime 72.0
Grok 41409
08Closedrank 130swe 75.0gpqa 87.5aime 95.0
09Closedrank 154gpqa 82.3aime 91.1
10Openrank 228gpqa 69.8

— not measured / not published · score has a caveat — hover the cell for details, see footnotes below.

Caveats

  • Qwen3 Max: all benchmark scores are vendor self-reported (Qwen), collected via the llm-stats aggregator — not independent measurements.
  • Grok 4 SWE-bench Verified: third-party aggregators only, medium reliability — xAI did not publish an official SWE-bench score.
  • GPT-5 mini LiveCodeBench v6 and MMLU-Pro: third-party measurement by LG AI (reasoning: high), not official OpenAI numbers.

Arena variants

  • Claude Sonnet 4.5 rated as claude-sonnet-4-5-20250929-high-32k. base variant: 1455 Elo, rank 58.
  • Claude Opus 4.1 rated as claude-opus-4-1-20250805-thinking-16k. base variant: 1447 Elo, rank 68.
  • GPT-5 rated as gpt-5-high (reasoning). gpt-5-chat variant: 1427 Elo, rank 98.
  • Qwen3 Max rated as qwen3-max-preview. only the preview variant is listed on the arena; final qwen3-max not found on the board.
  • DeepSeek V3.2 rated as deepseek-v3.2 (Thinking).
  • Gemini 2.5 Flash. preview-09-2025: 1404 Elo, rank 134.
  • Grok 4 rated as grok-4-0709.
  • GPT-5 mini rated as gpt-5-mini-high.
  • Llama 4 Maverick rated as llama-4-maverick-17b-128e-instruct. April 2025 listing showed 1417 Elo for an experimental chat version that did not match the released weights ('Leaderboard Illusion'). Current rating reflects the public weights..

The Llama 4 «Leaderboard Illusion»

April 2025 listing showed 1417 Elo for an experimental chat version that did not match the released weights ('Leaderboard Illusion'). Current rating reflects the public weights.

Sources & methodology

01

Arena ratings

LMArena Text Overall leaderboard snapshot: 7.78M blind pairwise votes across 392 models, snapshot Aug 13, 2026. Arena Elo measures human preference in chat, not benchmark accuracy — a model can top one board and lag on the other. arena leaderboard ↗

Benchmarks don't pay the bills — prices do. Compare what these models cost per 1M tokens.Browse API pricing →