Model Intelligence Ranking

Sorted by composite benchmark score
August 2026 · Data from BenchLM, Anthropic, OpenAI, DeepSeek

Benchmark Scores — Grouped Bar Chart

★ Wins (benchmarks topped)
Huihui-Qwen3.8-27B beats Claude Opus 4.6 Max on SWE-bench Pro & LiveCodeBench — while costing $0 to run.
It equals Fable 5 on LiveCodeBench (90.3% vs 90.0%) and beats GPT-5.5 on SWE-bench Pro by 3.1 points.
On terminal agent tasks it's 16 points ahead of DeepSeek V4 Flash — the other affordable open model.
Biggest gap: HLE (hardest reasoning). Frontier models are ~2× better, but you're running 27B parameters locally for free.