Benchmarks
Compare AI models
Pick two models to see release dates, benchmark scores, and community ratings side by side.
- Provider
- AnthropicOpenAIGoogle DeepMind
- Out yet?
- YesYesYes
- Status
- AvailableAvailableAvailable
- Release date
- Jul 1, 2026Apr 23, 2026Feb 19, 2026
- LM Arena?LM Arena's Elo-style rating from blind head-to-head votes: people compare two anonymous model answers and pick the better one. Higher is better.
- 1507.3106921725512Verify at LM ArenaAug 10, 20261476.6907176552236Verify at LM ArenaAug 10, 20261486.5107330227597Verify at LM ArenaAug 10, 2026
- Agentic?BenchLM.ai's score for multi-step agentic work - planning, tool use, and acting autonomously, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
- 75.34Verify at BenchLM.aiAug 10, 202659.67Verify at BenchLM.aiAug 10, 202677.2Verify at BenchLM.aiJul 14, 2026
- Coding?BenchLM.ai's score for code generation and software-engineering tasks, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
- 79.6Verify at BenchLM.aiAug 10, 202670.72Verify at BenchLM.aiAug 10, 202688.9Verify at BenchLM.aiJul 14, 2026
- InstructionFollowing?BenchLM.ai's score for following precise, detailed instructions, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
- 88.1Verify at BenchLM.aiJul 17, 202685.7Verify at BenchLM.aiJul 21, 202685.8Verify at BenchLM.aiJul 21, 2026
- Knowledge?BenchLM.ai's score for factual knowledge and question answering, normalized 0–100 across multiple benchmarks. Higher is better.
- 70.9Verify at BenchLM.aiAug 10, 202679.3Verify at BenchLM.aiAug 10, 202666.3Verify at BenchLM.aiAug 10, 2026
- Math?BenchLM.ai's score for mathematical problem solving, normalized 0–100 across multiple benchmarks. Higher is better.
- --61.6Verify at BenchLM.aiJul 17, 2026
- Multilingual?BenchLM.ai's score for capability across non-English languages, normalized 0–100 across multiple benchmarks. Higher is better.
- 100Verify at BenchLM.aiJul 17, 2026-100Verify at BenchLM.aiJul 17, 2026
- MultimodalGrounded?BenchLM.ai's score for understanding grounded in images and documents, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
- 79.7Verify at BenchLM.aiJul 17, 202668.2Verify at BenchLM.aiAug 10, 202680.6Verify at BenchLM.aiAug 10, 2026
- Reasoning?BenchLM.ai's score for logic and multi-step reasoning problems, normalized 0–100 across multiple benchmarks. Higher is better.
- --95.8Verify at BenchLM.aiJul 17, 2026
- ExploitBench?Epoch AI's ExploitBench: discovering and exploiting software vulnerabilities in controlled environments. Score is the percentage of tasks solved.Epoch AI
- -41.8Verify at Epoch AIAug 10, 202626.1Verify at Epoch AIAug 10, 2026
- FrontierCode Diamond?The hardest “Diamond” tier of Epoch AI's FrontierCode: research-level programming problems. Score is the percentage solved.Epoch AI
- -6.3Verify at Epoch AIJul 24, 20264.7Verify at Epoch AIJul 24, 2026
- GDP.pdf?Epoch AI's GDP.pdf benchmark: extracting and analyzing information from real-world PDF documents. Higher is better.Epoch AI
- 30Verify at Epoch AIAug 10, 202625Verify at Epoch AIAug 10, 202617Verify at Epoch AIAug 10, 2026
- Humanity's Last Exam?Humanity's Last Exam: expert-written questions across dozens of subjects, designed to sit far beyond what a web search can answer. Score is the percentage answered correctly.Epoch AI
- --46.4Verify at Epoch AIAug 10, 2026
- Vibe rating?OutYet's community rating: signed-in users score the model 1–10. Shown as the average and the number of votes.
- 9.7 / 10 · 3--
- Predecessor
- -
- Successor
- ---
| Open page → | Open page → | Open page → | |
|---|---|---|---|
| Provider | Anthropic | OpenAI | Google DeepMind |
| Out yet? | Yes | Yes | Yes |
| Status | Available | Available | Available |
| Release date | Jul 1, 2026 | Apr 23, 2026 | Feb 19, 2026 |
| LM Arena?LM Arena's Elo-style rating from blind head-to-head votes: people compare two anonymous model answers and pick the better one. Higher is better. | 1507.3106921725512Verify at LM ArenaAug 10, 2026 | 1476.6907176552236Verify at LM ArenaAug 10, 2026 | 1486.5107330227597Verify at LM ArenaAug 10, 2026 |
| Agentic?BenchLM.ai's score for multi-step agentic work - planning, tool use, and acting autonomously, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates | 75.34Verify at BenchLM.aiAug 10, 2026 | 59.67Verify at BenchLM.aiAug 10, 2026 | 77.2Verify at BenchLM.aiJul 14, 2026 |
| Coding?BenchLM.ai's score for code generation and software-engineering tasks, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates | 79.6Verify at BenchLM.aiAug 10, 2026 | 70.72Verify at BenchLM.aiAug 10, 2026 | 88.9Verify at BenchLM.aiJul 14, 2026 |
| InstructionFollowing?BenchLM.ai's score for following precise, detailed instructions, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates | 88.1Verify at BenchLM.aiJul 17, 2026 | 85.7Verify at BenchLM.aiJul 21, 2026 | 85.8Verify at BenchLM.aiJul 21, 2026 |
| Knowledge?BenchLM.ai's score for factual knowledge and question answering, normalized 0–100 across multiple benchmarks. Higher is better. | 70.9Verify at BenchLM.aiAug 10, 2026 | 79.3Verify at BenchLM.aiAug 10, 2026 | 66.3Verify at BenchLM.aiAug 10, 2026 |
| Math?BenchLM.ai's score for mathematical problem solving, normalized 0–100 across multiple benchmarks. Higher is better. | - | - | 61.6Verify at BenchLM.aiJul 17, 2026 |
| Multilingual?BenchLM.ai's score for capability across non-English languages, normalized 0–100 across multiple benchmarks. Higher is better. | 100Verify at BenchLM.aiJul 17, 2026 | - | 100Verify at BenchLM.aiJul 17, 2026 |
| MultimodalGrounded?BenchLM.ai's score for understanding grounded in images and documents, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates | 79.7Verify at BenchLM.aiJul 17, 2026 | 68.2Verify at BenchLM.aiAug 10, 2026 | 80.6Verify at BenchLM.aiAug 10, 2026 |
| Reasoning?BenchLM.ai's score for logic and multi-step reasoning problems, normalized 0–100 across multiple benchmarks. Higher is better. | - | - | 95.8Verify at BenchLM.aiJul 17, 2026 |
| ExploitBench?Epoch AI's ExploitBench: discovering and exploiting software vulnerabilities in controlled environments. Score is the percentage of tasks solved.Epoch AI | - | 41.8Verify at Epoch AIAug 10, 2026 | 26.1Verify at Epoch AIAug 10, 2026 |
| FrontierCode Diamond?The hardest “Diamond” tier of Epoch AI's FrontierCode: research-level programming problems. Score is the percentage solved.Epoch AI | - | 6.3Verify at Epoch AIJul 24, 2026 | 4.7Verify at Epoch AIJul 24, 2026 |
| GDP.pdf?Epoch AI's GDP.pdf benchmark: extracting and analyzing information from real-world PDF documents. Higher is better.Epoch AI | 30Verify at Epoch AIAug 10, 2026 | 25Verify at Epoch AIAug 10, 2026 | 17Verify at Epoch AIAug 10, 2026 |
| Humanity's Last Exam?Humanity's Last Exam: expert-written questions across dozens of subjects, designed to sit far beyond what a web search can answer. Score is the percentage answered correctly.Epoch AI | - | - | 46.4Verify at Epoch AIAug 10, 2026 |
| Vibe rating?OutYet's community rating: signed-in users score the model 1–10. Shown as the average and the number of votes. | 9.7 / 10 · 3 | - | - |
| Predecessor | - | GPT-5.4 | Gemini 3 Pro |
| Successor | - | - | - |
Benchmark scores are mirrored from third-party sources and captured on the dates shown. Numbers from different benchmarks, sources, or capture dates are not directly comparable.
Data from BenchLM.ai · Epoch AI, “AI Benchmarking Hub” (CC BY 4.0).