Eval · 2026-08-29 · 7 min
Verified is saturating. Pro is not.
29 August 2026. BenchLM Verified top three within a point. Scale live Pro and Terminal-Bench 3.0 still have air. Do not subtract them.
The live boards on 29 August 2026 are four objects. SWE-Verified (BenchLM) has Claude Opus 5 at 96.0, Mythos 5 at 95.5, Fable 5 at 95.0. The top three sit inside one point. That board is saturating. A model upgrade that moves Verified 0.4 points is a press release, not a procurement event.
Scale’s live Pro public (mini-swe-agent, uncapped, 250 turns) is a different object. Muse Spark 1.1* 61.5, gpt-5.4 xHigh* 59.1, Muse Spark* 55.0, claude-opus-4-6 thinking* 51.9. Paper-era Pro snapshots around 23% are not this board. We print both. We refuse to subtract them. The asterisks are the scaffold, not a footnote you can ignore.
Terminal-Bench 2.0 is also saturating: GPT-5.6 Sol 91.9, Mythos 5 88.0, Terra 87.4. Sol’s 91.9 is 88.8 as a solo agent and 91.9 when “ultra” coordinates four agents — a harness number printed as a model number. TB 2.1 (Artificial Analysis, Terminus 2, reward-hack zero) is tighter: Sol 89.5, Opus 5 89.1, Grok 4.6 88.4. TB 3.0 (74 tasks, 7 domains) is not saturating: Opus 5 42.7, Sol 34.6, Fable 5 34.0, Grok 4.5 15.7.
The desk pack is a fifth object. HM-SWE-Pro 400 holds the model and varies the shell. Loop+Claude 41.2, chat-without-a-harness 18.2. The landscape extension (Amp, Kiro, OpenHands, Aider, Junie, Cline, Goose, Qwen Code, Crush) is the same method, same window, labelled desk, conservative. It is not Scale live Pro and it is not Verified.
If you need one sentence for a steering committee: Verified no longer discriminates among frontier labs; Pro and TB 3.0 still do; harness swap still moves more than a model point on the boards that are not saturated. The eval desk is the table. This brief is the caption.