Benchmarks

Compare AI models

Pick two models to see release dates, benchmark scores, and community ratings side by side.

Open page →
Open page →
Open page →
Provider
Anthropic
OpenAI
Google DeepMind
Out yet?
Yes
Yes
Yes
Status
Available
Available
Available
Release date
Jul 1, 2026
Apr 23, 2026
Feb 19, 2026
LM Arena?LM Arena's Elo-style rating from blind head-to-head votes: people compare two anonymous model answers and pick the better one. Higher is better.Captured on different dates
1505.6827180827381Verify at LM ArenaSep 17, 2026
1476.084135658949Verify at LM ArenaSep 24, 2026
1486.809321925295Verify at LM ArenaSep 24, 2026
Agentic?BenchLM.ai's score for multi-step agentic work - planning, tool use, and acting autonomously, normalized 0–100 across multiple benchmarks. Higher is better.
73.8Verify at BenchLM.aiSep 24, 2026
60.21Verify at BenchLM.aiSep 24, 2026
38.92Verify at BenchLM.aiSep 24, 2026
Coding?BenchLM.ai's score for code generation and software-engineering tasks, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
74.56Verify at BenchLM.aiSep 24, 2026
64.54Verify at BenchLM.aiSep 24, 2026
88.9Verify at BenchLM.aiJul 14, 2026
InstructionFollowing?BenchLM.ai's score for following precise, detailed instructions, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
88.1Verify at BenchLM.aiJul 17, 2026
91.9Verify at BenchLM.aiSep 24, 2026
85.8Verify at BenchLM.aiJul 21, 2026
Knowledge?BenchLM.ai's score for factual knowledge and question answering, normalized 0–100 across multiple benchmarks. Higher is better.
82.59Verify at BenchLM.aiSep 24, 2026
70.59Verify at BenchLM.aiSep 24, 2026
64.19Verify at BenchLM.aiSep 24, 2026
Math?BenchLM.ai's score for mathematical problem solving, normalized 0–100 across multiple benchmarks. Higher is better.
-
-
61.6Verify at BenchLM.aiJul 17, 2026
Multilingual?BenchLM.ai's score for capability across non-English languages, normalized 0–100 across multiple benchmarks. Higher is better.
100Verify at BenchLM.aiJul 17, 2026
-
100Verify at BenchLM.aiJul 17, 2026
MultimodalGrounded?BenchLM.ai's score for understanding grounded in images and documents, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
79.7Verify at BenchLM.aiJul 17, 2026
71.4Verify at BenchLM.aiSep 24, 2026
79.1Verify at BenchLM.aiSep 24, 2026
Reasoning?BenchLM.ai's score for logic and multi-step reasoning problems, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
77.6Verify at BenchLM.aiSep 22, 2026
63.8Verify at BenchLM.aiSep 24, 2026
95.8Verify at BenchLM.aiJul 17, 2026
ExploitBench?Epoch AI's ExploitBench: discovering and exploiting software vulnerabilities in controlled environments. Score is the percentage of tasks solved.Epoch AI
-
41.8Verify at Epoch AISep 24, 2026
26.1Verify at Epoch AISep 24, 2026
FrontierCode Diamond?The hardest “Diamond” tier of Epoch AI's FrontierCode: research-level programming problems. Score is the percentage solved.Epoch AI
-
6.3Verify at Epoch AIJul 24, 2026
4.7Verify at Epoch AIJul 24, 2026
GDP.pdf?Epoch AI's GDP.pdf benchmark: extracting and analyzing information from real-world PDF documents. Higher is better.Epoch AI
30Verify at Epoch AISep 24, 2026
26Verify at Epoch AISep 24, 2026
17Verify at Epoch AISep 24, 2026
Humanity's Last Exam?Humanity's Last Exam: expert-written questions across dozens of subjects, designed to sit far beyond what a web search can answer. Score is the percentage answered correctly.Epoch AI
-
-
46.4Verify at Epoch AISep 24, 2026
Vibe rating?OutYet's community rating: signed-in users score the model 1–10. Shown as the average and the number of votes.
9.7 / 10 · 3
-
-
Successor
-
-
-

Benchmark scores are mirrored from third-party sources and captured on the dates shown. Numbers from different benchmarks, sources, or capture dates are not directly comparable.

Data from BenchLM.ai · Epoch AI, “AI Benchmarking Hub” (CC BY 4.0).