Huihui-Qwen3.8-27B beats Claude Opus 4.6 Max on SWE-bench Pro & LiveCodeBench — while costing $0 to run.
It equals Fable 5 on LiveCodeBench (90.3% vs 90.0%) and beats GPT-5.5 on SWE-bench Pro by 3.1 points.
On terminal agent tasks it's 16 points ahead of DeepSeek V4 Flash — the other affordable open model.
Biggest gap: HLE (hardest reasoning). Frontier models are ~2× better, but you're running 27B parameters locally for free.