HarnessmarketEnter the desk
/HMX Claude Code 29.1%/Open index 519 slips/Meridian desk 6,400 agents live/Skills + MCP 553,855 indexed (Skillful, Aug 2026)/Agent Skills spec 1.2M+ open packages/Deloitte: orchestration worth +15–30% of autonomous-agent TAM by 2030/Gartner: 40% of agentic projects cancelled by 2027 without a supervisor/SWE-bench Pro: harness swap > many model upgrades/HM-SWE-Pro: Loop+Claude 41.2 pass@1 · chat 18.2 · same model/Lewis 2608.26218: F2PF 28→49 under a tighter harness, same model/ACES: Skill Lift 0.21 · 947 paired cases · scan vs live ρ=0.14/Tokenomics: code review 59.4% of ChatDev tokens/HMX Claude Code 29.1%/Open index 519 slips/Meridian desk 6,400 agents live/Skills + MCP 553,855 indexed (Skillful, Aug 2026)/Agent Skills spec 1.2M+ open packages/Deloitte: orchestration worth +15–30% of autonomous-agent TAM by 2030/Gartner: 40% of agentic projects cancelled by 2027 without a supervisor/SWE-bench Pro: harness swap > many model upgrades/HM-SWE-Pro: Loop+Claude 41.2 pass@1 · chat 18.2 · same model/Lewis 2608.26218: F2PF 28→49 under a tighter harness, same model/ACES: Skill Lift 0.21 · 947 paired cases · scan vs live ρ=0.14/Tokenomics: code review 59.4% of ChatDev tokens

Eval · 2026-08-29 · 7 min

Verified is saturating. Pro is not.

29 August 2026. BenchLM Verified top three within a point. Scale live Pro and Terminal-Bench 3.0 still have air. Do not subtract them.

Harnessmarket Intelligence

The live boards on 29 August 2026 are four objects. SWE-Verified (BenchLM) has Claude Opus 5 at 96.0, Mythos 5 at 95.5, Fable 5 at 95.0. The top three sit inside one point. That board is saturating. A model upgrade that moves Verified 0.4 points is a press release, not a procurement event.

Scale’s live Pro public (mini-swe-agent, uncapped, 250 turns) is a different object. Muse Spark 1.1* 61.5, gpt-5.4 xHigh* 59.1, Muse Spark* 55.0, claude-opus-4-6 thinking* 51.9. Paper-era Pro snapshots around 23% are not this board. We print both. We refuse to subtract them. The asterisks are the scaffold, not a footnote you can ignore.

Terminal-Bench 2.0 is also saturating: GPT-5.6 Sol 91.9, Mythos 5 88.0, Terra 87.4. Sol’s 91.9 is 88.8 as a solo agent and 91.9 when “ultra” coordinates four agents — a harness number printed as a model number. TB 2.1 (Artificial Analysis, Terminus 2, reward-hack zero) is tighter: Sol 89.5, Opus 5 89.1, Grok 4.6 88.4. TB 3.0 (74 tasks, 7 domains) is not saturating: Opus 5 42.7, Sol 34.6, Fable 5 34.0, Grok 4.5 15.7.

The desk pack is a fifth object. HM-SWE-Pro 400 holds the model and varies the shell. Loop+Claude 41.2, chat-without-a-harness 18.2. The landscape extension (Amp, Kiro, OpenHands, Aider, Junie, Cline, Goose, Qwen Code, Crush) is the same method, same window, labelled desk, conservative. It is not Scale live Pro and it is not Verified.

If you need one sentence for a steering committee: Verified no longer discriminates among frontier labs; Pro and TB 3.0 still do; harness swap still moves more than a model point on the boards that are not saturated. The eval desk is the table. This brief is the caption.