Benchmarks

Compare AI models

Pick two models to see release dates, benchmark scores, and community ratings side by side.

Open page →
Open page →
Open page →
Provider
Anthropic
OpenAI
Google DeepMind
Out yet?
Yes
Yes
Yes
Status
Available
Available
Available
Release date
Jul 1, 2026
Apr 23, 2026
Feb 19, 2026
LM Arena?LM Arena's Elo-style rating from blind head-to-head votes: people compare two anonymous model answers and pick the better one. Higher is better.
1507.3106921725512Verify at LM ArenaAug 10, 2026
1476.6907176552236Verify at LM ArenaAug 10, 2026
1486.5107330227597Verify at LM ArenaAug 10, 2026
Agentic?BenchLM.ai's score for multi-step agentic work - planning, tool use, and acting autonomously, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
75.34Verify at BenchLM.aiAug 10, 2026
59.67Verify at BenchLM.aiAug 10, 2026
77.2Verify at BenchLM.aiJul 14, 2026
Coding?BenchLM.ai's score for code generation and software-engineering tasks, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
79.6Verify at BenchLM.aiAug 10, 2026
70.72Verify at BenchLM.aiAug 10, 2026
88.9Verify at BenchLM.aiJul 14, 2026
InstructionFollowing?BenchLM.ai's score for following precise, detailed instructions, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
88.1Verify at BenchLM.aiJul 17, 2026
85.7Verify at BenchLM.aiJul 21, 2026
85.8Verify at BenchLM.aiJul 21, 2026
Knowledge?BenchLM.ai's score for factual knowledge and question answering, normalized 0–100 across multiple benchmarks. Higher is better.
70.9Verify at BenchLM.aiAug 10, 2026
79.3Verify at BenchLM.aiAug 10, 2026
66.3Verify at BenchLM.aiAug 10, 2026
Math?BenchLM.ai's score for mathematical problem solving, normalized 0–100 across multiple benchmarks. Higher is better.
-
-
61.6Verify at BenchLM.aiJul 17, 2026
Multilingual?BenchLM.ai's score for capability across non-English languages, normalized 0–100 across multiple benchmarks. Higher is better.
100Verify at BenchLM.aiJul 17, 2026
-
100Verify at BenchLM.aiJul 17, 2026
MultimodalGrounded?BenchLM.ai's score for understanding grounded in images and documents, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
79.7Verify at BenchLM.aiJul 17, 2026
68.2Verify at BenchLM.aiAug 10, 2026
80.6Verify at BenchLM.aiAug 10, 2026
Reasoning?BenchLM.ai's score for logic and multi-step reasoning problems, normalized 0–100 across multiple benchmarks. Higher is better.
-
-
95.8Verify at BenchLM.aiJul 17, 2026
ExploitBench?Epoch AI's ExploitBench: discovering and exploiting software vulnerabilities in controlled environments. Score is the percentage of tasks solved.Epoch AI
-
41.8Verify at Epoch AIAug 10, 2026
26.1Verify at Epoch AIAug 10, 2026
FrontierCode Diamond?The hardest “Diamond” tier of Epoch AI's FrontierCode: research-level programming problems. Score is the percentage solved.Epoch AI
-
6.3Verify at Epoch AIJul 24, 2026
4.7Verify at Epoch AIJul 24, 2026
GDP.pdf?Epoch AI's GDP.pdf benchmark: extracting and analyzing information from real-world PDF documents. Higher is better.Epoch AI
30Verify at Epoch AIAug 10, 2026
25Verify at Epoch AIAug 10, 2026
17Verify at Epoch AIAug 10, 2026
Humanity's Last Exam?Humanity's Last Exam: expert-written questions across dozens of subjects, designed to sit far beyond what a web search can answer. Score is the percentage answered correctly.Epoch AI
-
-
46.4Verify at Epoch AIAug 10, 2026
Vibe rating?OutYet's community rating: signed-in users score the model 1–10. Shown as the average and the number of votes.
9.7 / 10 · 3
-
-
Successor
-
-
-

Benchmark scores are mirrored from third-party sources and captured on the dates shown. Numbers from different benchmarks, sources, or capture dates are not directly comparable.

Data from BenchLM.ai · Epoch AI, “AI Benchmarking Hub” (CC BY 4.0).