Benchmarks
Compare AI models Pick two models to see release dates, benchmark scores, and community ratings side by side.
Model 1 Choose a model… Claude Opus 5 Claude Fable 5 Claude Mythos 5 Claude Sonnet 5 Claude Opus 4.8 Claude Opus 4.7 Claude Sonnet 4.6 Claude Opus 4.6 Claude Opus 4.5 Claude Haiku 4.5 Claude Sonnet 4.5 Claude Opus 4.1 Claude Opus 4 Claude Sonnet 4 Claude 3.7 Sonnet Claude 3.5 Haiku Claude 3.5 Sonnet Claude 3 Haiku Claude 3 Opus Claude 3 Sonnet Claude 2.1 Claude 2 Claude 1 GPT-5.6 Luna GPT-5.6 Sol GPT-5.6 Terra GPT-Live GPT-5.4 GPT-5.5 Instant GPT-5.5 GPT-5.2 GPT-5.1 GPT-5 gpt-oss o3 o4-mini GPT-4.1 GPT-4.5 o3-mini o1 GPT-4o mini GPT-4o GPT-4 Turbo GPT-4 GPT-3.5 Turbo GPT-3 Nano Banana 2 Lite Gemini 3.5 Flash Gemini Omni Flash Gemini 3.1 Pro Gemini 3 Pro Gemini 2.5 Flash Gemini 2.5 Pro Gemma 3 Gemini 2.0 Flash Gemma 2 Gemini 1.5 Flash Gemma 1 Gemini 1.5 Pro Gemini 1.0 Muse Spark Llama 4 Llama 3.3 Llama 3.2 Llama 3.1 Llama 3 Llama 2 LLaMA Qwen 3.7 Max Qwen 3.6 Qwen 3.5 Qwen3-Coder-Next Qwen3-Max Qwen3-Coder Qwen 3 QwQ-32B Qwen 2.5 Qwen 2 Qwen 1.5 OLMo 3 OLMo 2 OLMo Amazon Nova 2 Lite Amazon Nova Premier Amazon Nova ERNIE 5.1 ERNIE 5.0 ERNIE 4.5 Seed-OSS 36B North Mini Code Command A Command R+ Command R DeepSeek V4 Flash DeepSeek V4 Pro DeepSeek-V3.2-Exp DeepSeek-V3.1 DeepSeek-R1 DeepSeek-V3 DeepSeek-V2.5 DeepSeek-V2 DeepSeek LLM Granite 4.1 Granite 4 Granite 3 Phi-4 Phi-3 MiniMax M3 MiniMax M2.7 MiniMax M2.5 MiniMax M2.1 MiniMax M2 MiniMax-M1 MiniMax-Text-01 Leanstral 1.5 Mistral Large 3 Magistral Mistral Medium 3 Mistral Small 3 Mistral Large 2 Mixtral 8x22B Mistral Large Mixtral 8x7B Mistral 7B Kimi K2.7 Code Kimi K2.6 Kimi K2.5 Kimi K2 Thinking Kimi K2 Kimi k1.5 Nemotron 3 Nemotron-4 340B Hunyuan-A13B Hunyuan TurboS Hunyuan-Large Grok 4.5 Grok 4.3 Grok 4.20 Grok 4.1 Grok 4 Heavy Grok 4 Grok 3 Grok-2 Grok-1.5 Grok-1 GLM-5.2 GLM-5.1 GLM-5 GLM-4.7 GLM-4.6 GLM-4.5 GLM-4 Open page → Model 2 Choose a model… Claude Opus 5 Claude Fable 5 Claude Mythos 5 Claude Sonnet 5 Claude Opus 4.8 Claude Opus 4.7 Claude Sonnet 4.6 Claude Opus 4.6 Claude Opus 4.5 Claude Haiku 4.5 Claude Sonnet 4.5 Claude Opus 4.1 Claude Opus 4 Claude Sonnet 4 Claude 3.7 Sonnet Claude 3.5 Haiku Claude 3.5 Sonnet Claude 3 Haiku Claude 3 Opus Claude 3 Sonnet Claude 2.1 Claude 2 Claude 1 GPT-5.6 Luna GPT-5.6 Sol GPT-5.6 Terra GPT-Live GPT-5.4 GPT-5.5 Instant GPT-5.5 GPT-5.2 GPT-5.1 GPT-5 gpt-oss o3 o4-mini GPT-4.1 GPT-4.5 o3-mini o1 GPT-4o mini GPT-4o GPT-4 Turbo GPT-4 GPT-3.5 Turbo GPT-3 Nano Banana 2 Lite Gemini 3.5 Flash Gemini Omni Flash Gemini 3.1 Pro Gemini 3 Pro Gemini 2.5 Flash Gemini 2.5 Pro Gemma 3 Gemini 2.0 Flash Gemma 2 Gemini 1.5 Flash Gemma 1 Gemini 1.5 Pro Gemini 1.0 Muse Spark Llama 4 Llama 3.3 Llama 3.2 Llama 3.1 Llama 3 Llama 2 LLaMA Qwen 3.7 Max Qwen 3.6 Qwen 3.5 Qwen3-Coder-Next Qwen3-Max Qwen3-Coder Qwen 3 QwQ-32B Qwen 2.5 Qwen 2 Qwen 1.5 OLMo 3 OLMo 2 OLMo Amazon Nova 2 Lite Amazon Nova Premier Amazon Nova ERNIE 5.1 ERNIE 5.0 ERNIE 4.5 Seed-OSS 36B North Mini Code Command A Command R+ Command R DeepSeek V4 Flash DeepSeek V4 Pro DeepSeek-V3.2-Exp DeepSeek-V3.1 DeepSeek-R1 DeepSeek-V3 DeepSeek-V2.5 DeepSeek-V2 DeepSeek LLM Granite 4.1 Granite 4 Granite 3 Phi-4 Phi-3 MiniMax M3 MiniMax M2.7 MiniMax M2.5 MiniMax M2.1 MiniMax M2 MiniMax-M1 MiniMax-Text-01 Leanstral 1.5 Mistral Large 3 Magistral Mistral Medium 3 Mistral Small 3 Mistral Large 2 Mixtral 8x22B Mistral Large Mixtral 8x7B Mistral 7B Kimi K2.7 Code Kimi K2.6 Kimi K2.5 Kimi K2 Thinking Kimi K2 Kimi k1.5 Nemotron 3 Nemotron-4 340B Hunyuan-A13B Hunyuan TurboS Hunyuan-Large Grok 4.5 Grok 4.3 Grok 4.20 Grok 4.1 Grok 4 Heavy Grok 4 Grok 3 Grok-2 Grok-1.5 Grok-1 GLM-5.2 GLM-5.1 GLM-5 GLM-4.7 GLM-4.6 GLM-4.5 GLM-4
Release date Jul 24, 2026
-
Agentic? BenchLM.ai's score for multi-step agentic work - planning, tool use, and acting autonomously, normalized 0–100 across multiple benchmarks. Higher is better. -
Coding? BenchLM.ai's score for code generation and software-engineering tasks, normalized 0–100 across multiple benchmarks. Higher is better. -
Knowledge? BenchLM.ai's score for factual knowledge and question answering, normalized 0–100 across multiple benchmarks. Higher is better. -
MultimodalGrounded? BenchLM.ai's score for understanding grounded in images and documents, normalized 0–100 across multiple benchmarks. Higher is better. -
Vibe rating? OutYet's community rating: signed-in users score the model 1–10. Shown as the average and the number of votes. 10.0 / 10 · 1
-
Open page → Model 1 Choose a model… Claude Opus 5 Claude Fable 5 Claude Mythos 5 Claude Sonnet 5 Claude Opus 4.8 Claude Opus 4.7 Claude Sonnet 4.6 Claude Opus 4.6 Claude Opus 4.5 Claude Haiku 4.5 Claude Sonnet 4.5 Claude Opus 4.1 Claude Opus 4 Claude Sonnet 4 Claude 3.7 Sonnet Claude 3.5 Haiku Claude 3.5 Sonnet Claude 3 Haiku Claude 3 Opus Claude 3 Sonnet Claude 2.1 Claude 2 Claude 1 GPT-5.6 Luna GPT-5.6 Sol GPT-5.6 Terra GPT-Live GPT-5.4 GPT-5.5 Instant GPT-5.5 GPT-5.2 GPT-5.1 GPT-5 gpt-oss o3 o4-mini GPT-4.1 GPT-4.5 o3-mini o1 GPT-4o mini GPT-4o GPT-4 Turbo GPT-4 GPT-3.5 Turbo GPT-3 Nano Banana 2 Lite Gemini 3.5 Flash Gemini Omni Flash Gemini 3.1 Pro Gemini 3 Pro Gemini 2.5 Flash Gemini 2.5 Pro Gemma 3 Gemini 2.0 Flash Gemma 2 Gemini 1.5 Flash Gemma 1 Gemini 1.5 Pro Gemini 1.0 Muse Spark Llama 4 Llama 3.3 Llama 3.2 Llama 3.1 Llama 3 Llama 2 LLaMA Qwen 3.7 Max Qwen 3.6 Qwen 3.5 Qwen3-Coder-Next Qwen3-Max Qwen3-Coder Qwen 3 QwQ-32B Qwen 2.5 Qwen 2 Qwen 1.5 OLMo 3 OLMo 2 OLMo Amazon Nova 2 Lite Amazon Nova Premier Amazon Nova ERNIE 5.1 ERNIE 5.0 ERNIE 4.5 Seed-OSS 36B North Mini Code Command A Command R+ Command R DeepSeek V4 Flash DeepSeek V4 Pro DeepSeek-V3.2-Exp DeepSeek-V3.1 DeepSeek-R1 DeepSeek-V3 DeepSeek-V2.5 DeepSeek-V2 DeepSeek LLM Granite 4.1 Granite 4 Granite 3 Phi-4 Phi-3 MiniMax M3 MiniMax M2.7 MiniMax M2.5 MiniMax M2.1 MiniMax M2 MiniMax-M1 MiniMax-Text-01 Leanstral 1.5 Mistral Large 3 Magistral Mistral Medium 3 Mistral Small 3 Mistral Large 2 Mixtral 8x22B Mistral Large Mixtral 8x7B Mistral 7B Kimi K2.7 Code Kimi K2.6 Kimi K2.5 Kimi K2 Thinking Kimi K2 Kimi k1.5 Nemotron 3 Nemotron-4 340B Hunyuan-A13B Hunyuan TurboS Hunyuan-Large Grok 4.5 Grok 4.3 Grok 4.20 Grok 4.1 Grok 4 Heavy Grok 4 Grok 3 Grok-2 Grok-1.5 Grok-1 GLM-5.2 GLM-5.1 GLM-5 GLM-4.7 GLM-4.6 GLM-4.5 GLM-4 Model 2 Choose a model… Claude Opus 5 Claude Fable 5 Claude Mythos 5 Claude Sonnet 5 Claude Opus 4.8 Claude Opus 4.7 Claude Sonnet 4.6 Claude Opus 4.6 Claude Opus 4.5 Claude Haiku 4.5 Claude Sonnet 4.5 Claude Opus 4.1 Claude Opus 4 Claude Sonnet 4 Claude 3.7 Sonnet Claude 3.5 Haiku Claude 3.5 Sonnet Claude 3 Haiku Claude 3 Opus Claude 3 Sonnet Claude 2.1 Claude 2 Claude 1 GPT-5.6 Luna GPT-5.6 Sol GPT-5.6 Terra GPT-Live GPT-5.4 GPT-5.5 Instant GPT-5.5 GPT-5.2 GPT-5.1 GPT-5 gpt-oss o3 o4-mini GPT-4.1 GPT-4.5 o3-mini o1 GPT-4o mini GPT-4o GPT-4 Turbo GPT-4 GPT-3.5 Turbo GPT-3 Nano Banana 2 Lite Gemini 3.5 Flash Gemini Omni Flash Gemini 3.1 Pro Gemini 3 Pro Gemini 2.5 Flash Gemini 2.5 Pro Gemma 3 Gemini 2.0 Flash Gemma 2 Gemini 1.5 Flash Gemma 1 Gemini 1.5 Pro Gemini 1.0 Muse Spark Llama 4 Llama 3.3 Llama 3.2 Llama 3.1 Llama 3 Llama 2 LLaMA Qwen 3.7 Max Qwen 3.6 Qwen 3.5 Qwen3-Coder-Next Qwen3-Max Qwen3-Coder Qwen 3 QwQ-32B Qwen 2.5 Qwen 2 Qwen 1.5 OLMo 3 OLMo 2 OLMo Amazon Nova 2 Lite Amazon Nova Premier Amazon Nova ERNIE 5.1 ERNIE 5.0 ERNIE 4.5 Seed-OSS 36B North Mini Code Command A Command R+ Command R DeepSeek V4 Flash DeepSeek V4 Pro DeepSeek-V3.2-Exp DeepSeek-V3.1 DeepSeek-R1 DeepSeek-V3 DeepSeek-V2.5 DeepSeek-V2 DeepSeek LLM Granite 4.1 Granite 4 Granite 3 Phi-4 Phi-3 MiniMax M3 MiniMax M2.7 MiniMax M2.5 MiniMax M2.1 MiniMax M2 MiniMax-M1 MiniMax-Text-01 Leanstral 1.5 Mistral Large 3 Magistral Mistral Medium 3 Mistral Small 3 Mistral Large 2 Mixtral 8x22B Mistral Large Mixtral 8x7B Mistral 7B Kimi K2.7 Code Kimi K2.6 Kimi K2.5 Kimi K2 Thinking Kimi K2 Kimi k1.5 Nemotron 3 Nemotron-4 340B Hunyuan-A13B Hunyuan TurboS Hunyuan-Large Grok 4.5 Grok 4.3 Grok 4.20 Grok 4.1 Grok 4 Heavy Grok 4 Grok 3 Grok-2 Grok-1.5 Grok-1 GLM-5.2 GLM-5.1 GLM-5 GLM-4.7 GLM-4.6 GLM-4.5 GLM-4 Provider Anthropic - Out yet? Yes - Status Available - Release date Jul 24, 2026 - Agentic? BenchLM.ai's score for multi-step agentic work - planning, tool use, and acting autonomously, normalized 0–100 across multiple benchmarks. Higher is better. 71.71Verify at BenchLM.ai Jul 28, 2026 - Coding? BenchLM.ai's score for code generation and software-engineering tasks, normalized 0–100 across multiple benchmarks. Higher is better. 77.77Verify at BenchLM.ai Jul 28, 2026 - Knowledge? BenchLM.ai's score for factual knowledge and question answering, normalized 0–100 across multiple benchmarks. Higher is better. 93.5Verify at BenchLM.ai Jul 28, 2026 - MultimodalGrounded? BenchLM.ai's score for understanding grounded in images and documents, normalized 0–100 across multiple benchmarks. Higher is better. 89.1Verify at BenchLM.ai Jul 28, 2026 - Vibe rating? OutYet's community rating: signed-in users score the model 1–10. Shown as the average and the number of votes. 10.0 / 10 · 1 - Predecessor - - Successor - -
Benchmark scores are mirrored from third-party sources and captured on the dates shown. Numbers from different benchmarks, sources, or capture dates are not directly comparable.
Data from BenchLM.ai .