Benchmarks

Compare AI models

Pick two models to see release dates, benchmark scores, and community ratings side by side.

Open page →
Provider
OpenAI
-
Out yet?
Yes
-
Status
Available
-
Release date
Jul 9, 2026
-
Agentic?BenchLM.ai's score for multi-step agentic work - planning, tool use, and acting autonomously, normalized 0–100 across multiple benchmarks. Higher is better.
68.79Verify at BenchLM.aiAug 20, 2026
-
Coding?BenchLM.ai's score for code generation and software-engineering tasks, normalized 0–100 across multiple benchmarks. Higher is better.
78.88Verify at BenchLM.aiAug 20, 2026
-
InstructionFollowing?BenchLM.ai's score for following precise, detailed instructions, normalized 0–100 across multiple benchmarks. Higher is better.
86.8Verify at BenchLM.aiJul 21, 2026
-
Knowledge?BenchLM.ai's score for factual knowledge and question answering, normalized 0–100 across multiple benchmarks. Higher is better.
83.5Verify at BenchLM.aiAug 20, 2026
-
MultimodalGrounded?BenchLM.ai's score for understanding grounded in images and documents, normalized 0–100 across multiple benchmarks. Higher is better.
84.8Verify at BenchLM.aiAug 20, 2026
-
GDP.pdf?Epoch AI's GDP.pdf benchmark: extracting and analyzing information from real-world PDF documents. Higher is better.Epoch AI
30.7Verify at Epoch AIAug 20, 2026
-
HealthBench?OpenAI's HealthBench evaluation of healthcare conversation quality. This value is vendor-reported; higher is better.openai-system-card
-
HealthBench Consensus?OpenAI's HealthBench Consensus evaluation. This value is vendor-reported; higher is better.openai-system-card
-
HealthBench Hard?The harder subset of OpenAI's HealthBench healthcare evaluation. This value is vendor-reported; higher is better.openai-system-card
-
HealthBench Professional?OpenAI's HealthBench Professional evaluation of model responses to realistic healthcare conversations. This value is vendor-reported; higher is better.openai-system-card
-
Vibe rating?OutYet's community rating: signed-in users score the model 1–10. Shown as the average and the number of votes.
-
-
Predecessor
-
-
Successor
-
-

Benchmark scores are mirrored from third-party sources and captured on the dates shown. Numbers from different benchmarks, sources, or capture dates are not directly comparable.

Data from BenchLM.ai · Epoch AI, “AI Benchmarking Hub” (CC BY 4.0).