How does a user manage to 'rate' the capability of AI models? There are many models currently available and considerable marketing hype about each. This blog entry, using AI aggregate scoring from multiple sources with AI Anthropic analysis, has rated the top 15 AI models.
What "capability" ranking measures: These composite scores blend several distinct test types into one number, and each type measures capabilities in a different manner:
- Broad knowledge tests — wide-ranging multiple-choice exams pulled from undergraduate and
graduate coursework across law, medicine, physics, history, and similar
subjects.
- Hard reasoning tests — questions written by subject-matter PhD graduates specifically to resist
being answered by a quick internet search. This is intended to isolate genuine
reasoning from memorized lookup.
- Real coding tests — the model is handed an actual error or 'bug' report from a genuine open-source
software project and has to produce a working solution, which is then checked
automatically against that project's own tests.
- Human preference voting — ordinary users are shown two anonymous model responses side by
side and vote for the one they prefer; the votes are aggregated into a
ranking similar to a chess rating system.
No model wins every category, and different trackers weight these tests differently when building a single composite score, so the exact position needs to be treated as approximate only, especially within the top cluster. There is no absolute answer nor position.
The Ranking of the top 15 as at September 2026
- Claude Opus 5 (Anthropic) — Tops the composite ranking at 63. Particularly strong on the hardreasoning tests and on real coding fixes/solutions.
- Claude Fable 5 (Anthropic) — Scores 62, close enough to Opus 5 that the gap plausibly reflects measurement noise rather than a real capability difference. Same underlying family as Opus 5, positioned as the lighter/more accessible counterpart.
- GPT-5.6 "Sol" (OpenAI) — Scores 61, tied with Grok 4.6. Strong across all four test categories rather than excelling in one; generally regarded as OpenAI's strongest all-purpose model as of mid-2026.
- Grok 4.6 (xAI) — Also scores 61. Notably strong in human preference voting specifically, meaning people rate its answers highly in direct side-by-side comparisons even where the formal test scores sit close to rivals.
- Gemini 3.1 Pro (Google) — Leads the broad knowledge test with 94.1% correct — a wide-ranging exam-style benchmark spanning many academic subjects. Strong generalist but trails the top cluster slightly on the hardest reasoning tests.
- GPT-5.5 (OpenAI) — OpenAI's prior flagship, since superseded internally by GPT-5.6, but still close to the frontier group.
- GLM-5.3 (Zhipu/Z.ai, China) — Scores 60, tied for the best-performing model whose underlying code and weights are published openly rather than kept proprietary. This means outside researchers and companies can download and run it themselves, rather than only accessing it through a paid, closed service.
- Kimi K3 (Moonshot AI, China) — Also scores 60, tied with GLM-5.3 as the strongest openly available model. Free to run for anyone with sufficient computing hardware, unlike the closed proprietary systems ranked above it.
- DeepSeek V4 (DeepSeek) — Openly available; strong reasoning and tool-use performance, slightly behind GLM-5.3 and Kimi K3 on the composite score.
- Qwen 3.6 (Alibaba) — Well suited to running locally on a user's own device rather than via a remote server; competitive coding performance.
- Llama 4 (Meta) — Solid general performance, weaker than the top group on the hardest reasoning tests.
- Mistral Large 3 / Devstral (Mistral AI) — Strong performance relative to its computing cost, particularly for coding tasks.
- Command A+ (Cohere) — Built for enterprise deployment; capable but not at the frontier.
- Ernie 5 (Baidu) — Chinese-developed; trails the top American labs on this composite ranking, though the gap has narrowed over 2026.
- Doubao 1.5 Pro (ByteDance) — Competent for consumer and search-oriented use; not benchmarked against the frontier test suites used above.



