Benchmarks
Compare AI models
Pick two models to see release dates, benchmark scores, and community ratings side by side.
- Provider
- AnthropicOpenAIGoogle DeepMind
- Out yet?
- YesYesYes
- Status
- AvailableAvailableAvailable
- Release date
- Jul 1, 2026Apr 23, 2026Feb 19, 2026
- LM Arena?LM Arena's Elo-style rating from blind head-to-head votes: people compare two anonymous model answers and pick the better one. Higher is better.Captured on different dates
- 1505.6827180827381Verify at LM ArenaSep 17, 20261476.084135658949Verify at LM ArenaSep 24, 20261486.809321925295Verify at LM ArenaSep 24, 2026
- Agentic?BenchLM.ai's score for multi-step agentic work - planning, tool use, and acting autonomously, normalized 0–100 across multiple benchmarks. Higher is better.
- 73.8Verify at BenchLM.aiSep 24, 202660.21Verify at BenchLM.aiSep 24, 202638.92Verify at BenchLM.aiSep 24, 2026
- Coding?BenchLM.ai's score for code generation and software-engineering tasks, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
- 74.56Verify at BenchLM.aiSep 24, 202664.54Verify at BenchLM.aiSep 24, 202688.9Verify at BenchLM.aiJul 14, 2026
- InstructionFollowing?BenchLM.ai's score for following precise, detailed instructions, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
- 88.1Verify at BenchLM.aiJul 17, 202691.9Verify at BenchLM.aiSep 24, 202685.8Verify at BenchLM.aiJul 21, 2026
- Knowledge?BenchLM.ai's score for factual knowledge and question answering, normalized 0–100 across multiple benchmarks. Higher is better.
- 82.59Verify at BenchLM.aiSep 24, 202670.59Verify at BenchLM.aiSep 24, 202664.19Verify at BenchLM.aiSep 24, 2026
- Math?BenchLM.ai's score for mathematical problem solving, normalized 0–100 across multiple benchmarks. Higher is better.
- --61.6Verify at BenchLM.aiJul 17, 2026
- Multilingual?BenchLM.ai's score for capability across non-English languages, normalized 0–100 across multiple benchmarks. Higher is better.
- 100Verify at BenchLM.aiJul 17, 2026-100Verify at BenchLM.aiJul 17, 2026
- MultimodalGrounded?BenchLM.ai's score for understanding grounded in images and documents, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
- 79.7Verify at BenchLM.aiJul 17, 202671.4Verify at BenchLM.aiSep 24, 202679.1Verify at BenchLM.aiSep 24, 2026
- Reasoning?BenchLM.ai's score for logic and multi-step reasoning problems, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates
- 77.6Verify at BenchLM.aiSep 22, 202663.8Verify at BenchLM.aiSep 24, 202695.8Verify at BenchLM.aiJul 17, 2026
- ExploitBench?Epoch AI's ExploitBench: discovering and exploiting software vulnerabilities in controlled environments. Score is the percentage of tasks solved.Epoch AI
- -41.8Verify at Epoch AISep 24, 202626.1Verify at Epoch AISep 24, 2026
- FrontierCode Diamond?The hardest “Diamond” tier of Epoch AI's FrontierCode: research-level programming problems. Score is the percentage solved.Epoch AI
- -6.3Verify at Epoch AIJul 24, 20264.7Verify at Epoch AIJul 24, 2026
- GDP.pdf?Epoch AI's GDP.pdf benchmark: extracting and analyzing information from real-world PDF documents. Higher is better.Epoch AI
- 30Verify at Epoch AISep 24, 202626Verify at Epoch AISep 24, 202617Verify at Epoch AISep 24, 2026
- Humanity's Last Exam?Humanity's Last Exam: expert-written questions across dozens of subjects, designed to sit far beyond what a web search can answer. Score is the percentage answered correctly.Epoch AI
- --46.4Verify at Epoch AISep 24, 2026
- Vibe rating?OutYet's community rating: signed-in users score the model 1–10. Shown as the average and the number of votes.
- 9.7 / 10 · 3--
- Predecessor
- -
- Successor
- ---
| Open page → | Open page → | Open page → | |
|---|---|---|---|
| Provider | Anthropic | OpenAI | Google DeepMind |
| Out yet? | Yes | Yes | Yes |
| Status | Available | Available | Available |
| Release date | Jul 1, 2026 | Apr 23, 2026 | Feb 19, 2026 |
| LM Arena?LM Arena's Elo-style rating from blind head-to-head votes: people compare two anonymous model answers and pick the better one. Higher is better.Captured on different dates | 1505.6827180827381Verify at LM ArenaSep 17, 2026 | 1476.084135658949Verify at LM ArenaSep 24, 2026 | 1486.809321925295Verify at LM ArenaSep 24, 2026 |
| Agentic?BenchLM.ai's score for multi-step agentic work - planning, tool use, and acting autonomously, normalized 0–100 across multiple benchmarks. Higher is better. | 73.8Verify at BenchLM.aiSep 24, 2026 | 60.21Verify at BenchLM.aiSep 24, 2026 | 38.92Verify at BenchLM.aiSep 24, 2026 |
| Coding?BenchLM.ai's score for code generation and software-engineering tasks, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates | 74.56Verify at BenchLM.aiSep 24, 2026 | 64.54Verify at BenchLM.aiSep 24, 2026 | 88.9Verify at BenchLM.aiJul 14, 2026 |
| InstructionFollowing?BenchLM.ai's score for following precise, detailed instructions, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates | 88.1Verify at BenchLM.aiJul 17, 2026 | 91.9Verify at BenchLM.aiSep 24, 2026 | 85.8Verify at BenchLM.aiJul 21, 2026 |
| Knowledge?BenchLM.ai's score for factual knowledge and question answering, normalized 0–100 across multiple benchmarks. Higher is better. | 82.59Verify at BenchLM.aiSep 24, 2026 | 70.59Verify at BenchLM.aiSep 24, 2026 | 64.19Verify at BenchLM.aiSep 24, 2026 |
| Math?BenchLM.ai's score for mathematical problem solving, normalized 0–100 across multiple benchmarks. Higher is better. | - | - | 61.6Verify at BenchLM.aiJul 17, 2026 |
| Multilingual?BenchLM.ai's score for capability across non-English languages, normalized 0–100 across multiple benchmarks. Higher is better. | 100Verify at BenchLM.aiJul 17, 2026 | - | 100Verify at BenchLM.aiJul 17, 2026 |
| MultimodalGrounded?BenchLM.ai's score for understanding grounded in images and documents, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates | 79.7Verify at BenchLM.aiJul 17, 2026 | 71.4Verify at BenchLM.aiSep 24, 2026 | 79.1Verify at BenchLM.aiSep 24, 2026 |
| Reasoning?BenchLM.ai's score for logic and multi-step reasoning problems, normalized 0–100 across multiple benchmarks. Higher is better.Captured on different dates | 77.6Verify at BenchLM.aiSep 22, 2026 | 63.8Verify at BenchLM.aiSep 24, 2026 | 95.8Verify at BenchLM.aiJul 17, 2026 |
| ExploitBench?Epoch AI's ExploitBench: discovering and exploiting software vulnerabilities in controlled environments. Score is the percentage of tasks solved.Epoch AI | - | 41.8Verify at Epoch AISep 24, 2026 | 26.1Verify at Epoch AISep 24, 2026 |
| FrontierCode Diamond?The hardest “Diamond” tier of Epoch AI's FrontierCode: research-level programming problems. Score is the percentage solved.Epoch AI | - | 6.3Verify at Epoch AIJul 24, 2026 | 4.7Verify at Epoch AIJul 24, 2026 |
| GDP.pdf?Epoch AI's GDP.pdf benchmark: extracting and analyzing information from real-world PDF documents. Higher is better.Epoch AI | 30Verify at Epoch AISep 24, 2026 | 26Verify at Epoch AISep 24, 2026 | 17Verify at Epoch AISep 24, 2026 |
| Humanity's Last Exam?Humanity's Last Exam: expert-written questions across dozens of subjects, designed to sit far beyond what a web search can answer. Score is the percentage answered correctly.Epoch AI | - | - | 46.4Verify at Epoch AISep 24, 2026 |
| Vibe rating?OutYet's community rating: signed-in users score the model 1–10. Shown as the average and the number of votes. | 9.7 / 10 · 3 | - | - |
| Predecessor | - | GPT-5.4 | Gemini 3 Pro |
| Successor | - | - | - |
Benchmark scores are mirrored from third-party sources and captured on the dates shown. Numbers from different benchmarks, sources, or capture dates are not directly comparable.
Data from BenchLM.ai · Epoch AI, “AI Benchmarking Hub” (CC BY 4.0).