the Ledger
LLM Benchmark Comparison
Compare performance of frontier AI models across industry-standard benchmarks. Our benchmark data includes verified scores from official model releases and research papers, covering capabilities like general knowledge (MMLU), code generation (HumanEval), mathematical reasoning (GSM8K), and more.
View benchmark methodology and glossary
Keyboard Shortcuts:
- / - Focus search input
- Esc - Clear all filters
What capability costs you
Curated models with at least 2 reviewed reasoning/coding results, both reviewed price columns, and a cost above zero — plotted against what they would cost you: not per million tokens, but per month at 30 prompts a day, about 1,500 tokens of context in and 600 tokens out on each — one person using a model steadily, not an application serving traffic. Anything that misses one of those is named underneath, with the reason.
- $9.18 a month · $0.31 a day
- 68th percentile across 3 benchmarks
- Matches a $20/month plan at 65 prompts a day
- $2.00 in / $12.00 out per 1M tokens
- $11.47 a month · $0.38 a day
- 30th percentile across 4 benchmarks
- Matches a $20/month plan at 52 prompts a day
- $2.50 in / $15.00 out per 1M tokens
- $20.25 a month · $0.67 a day
- 92nd percentile across 2 benchmarks
- Matches a $20/month plan at 30 prompts a day
- $5.00 in / $25.00 out per 1M tokens
- $20.25 a month · $0.67 a day
- 33rd percentile across 2 benchmarks
- Matches a $20/month plan at 30 prompts a day
- $5.00 in / $25.00 out per 1M tokens
- $22.95 a month · $0.76 a day
- 69th percentile across 3 benchmarks
- Matches a $20/month plan at 26 prompts a day
- $5.00 in / $30.00 out per 1M tokens
- $138 a month · $4.59 a day
- 77th percentile across 2 benchmarks
- Matches a $20/month plan at 4.4 prompts a day
- $30.00 in / $180.00 out per 1M tokens
Showing the 28 curated models of 496 tracked — reviewed, sourced and dated, ranked by performance across benchmarks. The rest are unreviewed scrapes, kept in the full view.
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1493 ELO✓
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 5 $/1M tokens✓
- Output Cost ($/1M tokens)
- 25 $/1M tokens✓
- SWE-bench Verified (%)
- 83.5%✓
- Terminal-Bench 2.0 (%)
- 90.2%✓
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- GPQA Diamond (%)
- 94.6%✓
- Humanity's Last Exam (%)
- 44.3%✓
- Input Cost ($/1M tokens)
- 30 $/1M tokens✓
- Output Cost ($/1M tokens)
- 180 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1477 ELO✓
- GPQA Diamond (%)
- 94%✓
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 5 $/1M tokens✓
- Output Cost ($/1M tokens)
- 30 $/1M tokens✓
- SWE-bench Verified (%)
- 80.6%✓
- Terminal-Bench 2.0 (%)
- 84.7%✓
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1487 ELO✓
- GPQA Diamond (%)
- 94.1%✓
- Humanity's Last Exam (%)
- 46.4%✓
- Input Cost ($/1M tokens)
- 2 $/1M tokens✓
- Output Cost ($/1M tokens)
- 12 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- 80.2%✓
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1497 ELO✓
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 5 $/1M tokens✓
- Output Cost ($/1M tokens)
- 25 $/1M tokens✓
- SWE-bench Verified (%)
- 78.7%✓
- Terminal-Bench 2.0 (%)
- 79.8%✓
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1477 ELO✓
- GPQA Diamond (%)
- 93.3%✓
- Humanity's Last Exam (%)
- 36.2%✓
- Input Cost ($/1M tokens)
- 2.5 $/1M tokens✓
- Output Cost ($/1M tokens)
- 15 $/1M tokens✓
- SWE-bench Verified (%)
- 76.9%✓
- Terminal-Bench 2.0 (%)
- 81.8%✓
- ARC-AGI-1 (accuracy)
- 75 accuracy
- ARC-AGI-2 (accuracy)
- 31.11 accuracy
- Arena ELO
- 1486 ELO✓
- GPQA Diamond (%)
- 91.9%✓
- Humanity's Last Exam (%)
- 37.5%✓
- Input Cost ($/1M tokens)
- Output Cost ($/1M tokens)
- SWE-bench Verified (%)
- 76.2%✓
- Terminal-Bench 2.0 (%)
- 54.2%✓
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1476 ELO✓
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 1.5 $/1M tokens✓
- Output Cost ($/1M tokens)
- 9 $/1M tokens✓
- SWE-bench Verified (%)
- 79.3%✓
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- GPQA Diamond (%)
- 93.9%✓
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 30 $/1M tokens✓
- Output Cost ($/1M tokens)
- 180 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1507 ELO✓
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 10 $/1M tokens✓
- Output Cost ($/1M tokens)
- 50 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- 14.33 accuracy
- ARC-AGI-2 (accuracy)
- 1.25 accuracy
- Arena ELO
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 1 $/1M tokens✓
- Output Cost ($/1M tokens)
- 5 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 10 $/1M tokens✓
- Output Cost ($/1M tokens)
- 50 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1473 ELO✓
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 5 $/1M tokens✓
- Output Cost ($/1M tokens)
- 25 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1493 ELO✓
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 5 $/1M tokens✓
- Output Cost ($/1M tokens)
- 25 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1472 ELO✓
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 3 $/1M tokens✓
- Output Cost ($/1M tokens)
- 15 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 2 $/1M tokens✓
- Output Cost ($/1M tokens)
- 10 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 0.14 $/1M tokens✓
- Output Cost ($/1M tokens)
- 0.28 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 0.43 $/1M tokens✓
- Output Cost ($/1M tokens)
- 0.87 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 1.25 $/1M tokens✓
- Output Cost ($/1M tokens)
- 10 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1485 ELO✓
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 1.5 $/1M tokens✓
- Output Cost ($/1M tokens)
- 7.5 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 0.2 $/1M tokens✓
- Output Cost ($/1M tokens)
- 1.2 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1482 ELO✓
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 5 $/1M tokens✓
- Output Cost ($/1M tokens)
- 30 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 2 $/1M tokens✓
- Output Cost ($/1M tokens)
- 12 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 1.25 $/1M tokens✓
- Output Cost ($/1M tokens)
- 2.5 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- 2 $/1M tokens✓
- Output Cost ($/1M tokens)
- 6 $/1M tokens✓
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1485 ELO✓
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- Output Cost ($/1M tokens)
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1475 ELO✓
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- Output Cost ($/1M tokens)
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
- ARC-AGI-1 (accuracy)
- ARC-AGI-2 (accuracy)
- Arena ELO
- 1497 ELO✓
- GPQA Diamond (%)
- Humanity's Last Exam (%)
- Input Cost ($/1M tokens)
- Output Cost ($/1M tokens)
- SWE-bench Verified (%)
- Terminal-Bench 2.0 (%)
the pulse
What moved on the board
Changes to the curated tier over the last 60 days — models added, scores corrected, positions on the frontier changing hands. Each one landed as a reviewed diff, so the date is the date it reached the board.
The curated board opened with 28 flagship models
The Ledger's default view stopped being a 471-model dump. These models are scored on the headline benchmark set, and every number arrives as a reviewed diff carrying its own source URL, test date and named reviewer.
Methodology & Glossary
Benchmark methodology and detailed glossary coming soon.