Benchmarks

Compare AI models

Pick two models to see release dates, benchmark scores, and community ratings side by side.

Open page →
Provider
Alibaba Qwen
-
Out yet?
Yes
-
Status
Available
-
Release date
Sep 19, 2024
-
Agentic?BenchLM.ai's score for multi-step agentic work - planning, tool use, and acting autonomously, normalized 0–100 across multiple benchmarks. Higher is better.
45.8Verify at BenchLM.aiJul 3, 2026
-
Math?BenchLM.ai's score for mathematical problem solving, normalized 0–100 across multiple benchmarks. Higher is better.
76.2Verify at BenchLM.aiJul 3, 2026
-
Multilingual?BenchLM.ai's score for capability across non-English languages, normalized 0–100 across multiple benchmarks. Higher is better.
61.6Verify at BenchLM.aiJul 17, 2026
-
Reasoning?BenchLM.ai's score for logic and multi-step reasoning problems, normalized 0–100 across multiple benchmarks. Higher is better.
56.6Verify at BenchLM.aiJul 17, 2026
-
Vibe rating?OutYet's community rating: signed-in users score the model 1–10. Shown as the average and the number of votes.
-
-
Predecessor
-
Successor
-

Benchmark scores are mirrored from third-party sources and captured on the dates shown. Numbers from different benchmarks, sources, or capture dates are not directly comparable.

Data from BenchLM.ai.