Skip to main content
wideriver_

the Ledger

LLM Benchmark Comparison

Compare performance of frontier AI models across industry-standard benchmarks. Our benchmark data includes verified scores from official model releases and research papers, covering capabilities like general knowledge (MMLU), code generation (HumanEval), mathematical reasoning (GSM8K), and more.

Last updated: September 9, 2026

View benchmark methodology and glossary

Keyboard Shortcuts:

  • / - Focus search input
  • Esc - Clear all filters

What capability costs you

Curated models with at least 2 reviewed reasoning/coding results, both reviewed price columns, and a cost above zero — plotted against what they would cost you: not per million tokens, but per month at 30 prompts a day, about 1,500 tokens of context in and 600 tokens out on each — one person using a model steadily, not an application serving traffic. Anything that misses one of those is named underneath, with the reason.

Capability is the same mean percentile the board ranks by, normalized within each benchmark so results on different tests compare — SWE-bench Verified, GPQA Diamond, Terminal-Bench 2.0, Humanity's Last Exam. Cost is derived from the reviewed Input and Output Cost columns — 6 models plotted, 2 on the value frontier. The $20/month plan figure is a reference price (the price point the mainstream consumer chat plans sit at — check it against whatever you actually pay), not a quote.

  1. Gemini 3.1 Pro PreviewBest value at its price

    Google

    Cost
    $9.18 a month · $0.31 a day
    Capability
    68th percentile across 3 benchmarks
    Break-even
    Matches a $20/month plan at 65 prompts a day
    List price
    $2.00 in / $12.00 out per 1M tokens
  2. GPT-5.4Beaten on value

    OpenAI

    Cost
    $11.47 a month · $0.38 a day
    Capability
    30th percentile across 4 benchmarks
    Break-even
    Matches a $20/month plan at 52 prompts a day
    List price
    $2.50 in / $15.00 out per 1M tokens
  3. Claude Opus 4.7Best value at its price

    Anthropic

    Cost
    $20.25 a month · $0.67 a day
    Capability
    92nd percentile across 2 benchmarks
    Break-even
    Matches a $20/month plan at 30 prompts a day
    List price
    $5.00 in / $25.00 out per 1M tokens
  4. Claude Opus 4.6Beaten on value

    Anthropic

    Cost
    $20.25 a month · $0.67 a day
    Capability
    33rd percentile across 2 benchmarks
    Break-even
    Matches a $20/month plan at 30 prompts a day
    List price
    $5.00 in / $25.00 out per 1M tokens
  5. GPT-5.5Beaten on value

    OpenAI

    Cost
    $22.95 a month · $0.76 a day
    Capability
    69th percentile across 3 benchmarks
    Break-even
    Matches a $20/month plan at 26 prompts a day
    List price
    $5.00 in / $30.00 out per 1M tokens
  6. GPT-5.4 ProBeaten on value

    OpenAI

    Cost
    $138 a month · $4.59 a day
    Capability
    77th percentile across 2 benchmarks
    Break-even
    Matches a $20/month plan at 4.4 prompts a day
    List price
    $30.00 in / $180.00 out per 1M tokens

Not plotted — no capability score in the curated set yet: Claude Fable 5, Claude Haiku 4.5, Claude Mythos 5, Claude Opus 4.8, Claude Opus 5, Claude Sonnet 4.6, Claude Sonnet 5, DeepSeek-V4-Flash, DeepSeek-V4-Pro, Gemini 2.5 Pro, Gemini 3.6 Flash, GPT-5.6 Luna, GPT-5.6 Sol, GPT-5.6 Terra, Grok 4.3, Grok 4.5, Kimi K3, Qwen3.7 Max, Qwen3.8 MaxNot plotted — fewer than two reviewed reasoning/coding results in the curated set yet: Gemini 3.5 Flash, GPT-5.5 ProNot plotted — no published price in the curated set yet: Gemini 3 Pro

Showing the 28 curated models of 496 tracked — reviewed, sourced and dated, ranked by performance across benchmarks. The rest are unreviewed scrapes, kept in the full view.

28 models found
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1493 ELO
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    5 $/1M tokens
    Output Cost ($/1M tokens)
    25 $/1M tokens
    SWE-bench Verified (%)
    83.5%
    Terminal-Bench 2.0 (%)
    90.2%
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    Not tested
    GPQA Diamond (%)
    94.6%
    Humanity's Last Exam (%)
    44.3%
    Input Cost ($/1M tokens)
    30 $/1M tokens
    Output Cost ($/1M tokens)
    180 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • GPT-5.5

    OpenAI

    ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1477 ELO
    GPQA Diamond (%)
    94%
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    5 $/1M tokens
    Output Cost ($/1M tokens)
    30 $/1M tokens
    SWE-bench Verified (%)
    80.6%
    Terminal-Bench 2.0 (%)
    84.7%
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1487 ELO
    GPQA Diamond (%)
    94.1%
    Humanity's Last Exam (%)
    46.4%
    Input Cost ($/1M tokens)
    2 $/1M tokens
    Output Cost ($/1M tokens)
    12 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    80.2%
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1497 ELO
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    5 $/1M tokens
    Output Cost ($/1M tokens)
    25 $/1M tokens
    SWE-bench Verified (%)
    78.7%
    Terminal-Bench 2.0 (%)
    79.8%
  • GPT-5.4

    OpenAI

    ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1477 ELO
    GPQA Diamond (%)
    93.3%
    Humanity's Last Exam (%)
    36.2%
    Input Cost ($/1M tokens)
    2.5 $/1M tokens
    Output Cost ($/1M tokens)
    15 $/1M tokens
    SWE-bench Verified (%)
    76.9%
    Terminal-Bench 2.0 (%)
    81.8%
  • ARC-AGI-1 (accuracy)
    75 accuracy
    ARC-AGI-2 (accuracy)
    31.11 accuracy
    Arena ELO
    1486 ELO
    GPQA Diamond (%)
    91.9%
    Humanity's Last Exam (%)
    37.5%
    Input Cost ($/1M tokens)
    Not tested
    Output Cost ($/1M tokens)
    Not tested
    SWE-bench Verified (%)
    76.2%
    Terminal-Bench 2.0 (%)
    54.2%
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1476 ELO
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    1.5 $/1M tokens
    Output Cost ($/1M tokens)
    9 $/1M tokens
    SWE-bench Verified (%)
    79.3%
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    Not tested
    GPQA Diamond (%)
    93.9%
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    30 $/1M tokens
    Output Cost ($/1M tokens)
    180 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1507 ELO
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    10 $/1M tokens
    Output Cost ($/1M tokens)
    50 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    14.33 accuracy
    ARC-AGI-2 (accuracy)
    1.25 accuracy
    Arena ELO
    Not tested
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    1 $/1M tokens
    Output Cost ($/1M tokens)
    5 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    Not tested
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    10 $/1M tokens
    Output Cost ($/1M tokens)
    50 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1473 ELO
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    5 $/1M tokens
    Output Cost ($/1M tokens)
    25 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • Claude Opus 5

    Anthropic

    ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1493 ELO
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    5 $/1M tokens
    Output Cost ($/1M tokens)
    25 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1472 ELO
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    3 $/1M tokens
    Output Cost ($/1M tokens)
    15 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    Not tested
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    2 $/1M tokens
    Output Cost ($/1M tokens)
    10 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    Not tested
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    0.14 $/1M tokens
    Output Cost ($/1M tokens)
    0.28 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    Not tested
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    0.43 $/1M tokens
    Output Cost ($/1M tokens)
    0.87 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    Not tested
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    1.25 $/1M tokens
    Output Cost ($/1M tokens)
    10 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1485 ELO
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    1.5 $/1M tokens
    Output Cost ($/1M tokens)
    7.5 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    Not tested
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    0.2 $/1M tokens
    Output Cost ($/1M tokens)
    1.2 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1482 ELO
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    5 $/1M tokens
    Output Cost ($/1M tokens)
    30 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    Not tested
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    2 $/1M tokens
    Output Cost ($/1M tokens)
    12 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    Not tested
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    1.25 $/1M tokens
    Output Cost ($/1M tokens)
    2.5 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    Not tested
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    2 $/1M tokens
    Output Cost ($/1M tokens)
    6 $/1M tokens
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • Kimi K3

    Moonshot AI

    ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1485 ELO
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    Not tested
    Output Cost ($/1M tokens)
    Not tested
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1475 ELO
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    Not tested
    Output Cost ($/1M tokens)
    Not tested
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested
  • ARC-AGI-1 (accuracy)
    Not tested
    ARC-AGI-2 (accuracy)
    Not tested
    Arena ELO
    1497 ELO
    GPQA Diamond (%)
    Not tested
    Humanity's Last Exam (%)
    Not tested
    Input Cost ($/1M tokens)
    Not tested
    Output Cost ($/1M tokens)
    Not tested
    SWE-bench Verified (%)
    Not tested
    Terminal-Bench 2.0 (%)
    Not tested

Method: curated scores are sourced, dated and reviewed · marks a reviewed score · unreviewed rows are scraped and labelled as such

the pulse

What moved on the board

Changes to the curated tier over the last 60 days — models added, scores corrected, positions on the frontier changing hands. Each one landed as a reviewed diff, so the date is the date it reached the board.

  1. new entries

    The curated board opened with 28 flagship models

    The Ledger's default view stopped being a 471-model dump. These models are scored on the headline benchmark set, and every number arrives as a reviewed diff carrying its own source URL, test date and named reviewer.

    Claude Fable 5, Claude Haiku 4.5, Claude Mythos 5, Claude Opus 4.6, Claude Opus 4.7, Claude Opus 4.8, Claude Opus 5, Claude Sonnet 4.6 and 20 more

    Reviewed by agent-opsThe change

Board changes that qualify as news also run on the Wire.Full wire →

Methodology & Glossary

Benchmark methodology and detailed glossary coming soon.