Benchmarks

Compare AI models

Pick two models to see release dates, benchmark scores, and community ratings side by side.

Open page →
Provider
Moonshot AI
-
Out yet?
Yes
-
Status
Available
-
Release date
Jan 27, 2026
-
Agentic?BenchLM.ai's score for multi-step agentic work - planning, tool use, and acting autonomously, normalized 0–100 across multiple benchmarks. Higher is better.
49.5Verify at BenchLM.aiJul 14, 2026
-
Coding?BenchLM.ai's score for code generation and software-engineering tasks, normalized 0–100 across multiple benchmarks. Higher is better.
54.28Verify at BenchLM.aiAug 16, 2026
-
InstructionFollowing?BenchLM.ai's score for following precise, detailed instructions, normalized 0–100 across multiple benchmarks. Higher is better.
91Verify at BenchLM.aiAug 16, 2026
-
Knowledge?BenchLM.ai's score for factual knowledge and question answering, normalized 0–100 across multiple benchmarks. Higher is better.
57.4Verify at BenchLM.aiAug 16, 2026
-
Math?BenchLM.ai's score for mathematical problem solving, normalized 0–100 across multiple benchmarks. Higher is better.
62.5Verify at BenchLM.aiAug 16, 2026
-
Multilingual?BenchLM.ai's score for capability across non-English languages, normalized 0–100 across multiple benchmarks. Higher is better.
38.2Verify at BenchLM.aiAug 16, 2026
-
MultimodalGrounded?BenchLM.ai's score for understanding grounded in images and documents, normalized 0–100 across multiple benchmarks. Higher is better.
62.5Verify at BenchLM.aiJul 17, 2026
-
Reasoning?BenchLM.ai's score for logic and multi-step reasoning problems, normalized 0–100 across multiple benchmarks. Higher is better.
51Verify at BenchLM.aiJul 17, 2026
-
FrontierCode Diamond?The hardest “Diamond” tier of Epoch AI's FrontierCode: research-level programming problems. Score is the percentage solved.Epoch AI
1Verify at Epoch AIJul 24, 2026
-
Humanity's Last Exam?Humanity's Last Exam: expert-written questions across dozens of subjects, designed to sit far beyond what a web search can answer. Score is the percentage answered correctly.Epoch AI
24.4Verify at Epoch AIAug 16, 2026
-
SWE-Bench Pro?SWE-Bench Pro: resolving real GitHub issues in large codebases end to end. Score is the percentage of issues resolved.Hugging Face
50.7Verify at Hugging FaceAug 16, 2026
-
Vibe rating?OutYet's community rating: signed-in users score the model 1–10. Shown as the average and the number of votes.
-
-
Predecessor
-
Successor
-

Benchmark scores are mirrored from third-party sources and captured on the dates shown. Numbers from different benchmarks, sources, or capture dates are not directly comparable.

Data from BenchLM.ai · Epoch AI, “AI Benchmarking Hub” (CC BY 4.0).