Eval desk · through 2026-08-29
Hold the model. The harness is the independent variable.
On SWE-bench Pro, swapping the harness moved pass@1 more than many model upgrades. The literature now says the same thing in a dozen independent runs: 10–36 points of harness, Skill Lift near 0.21, review eating 59% of tokens, MCP success still under 45% at the frontier. Desk pack plus sources — labelled, not blended.
HM-SWE-Pro Δ
+22.4 pt
Best harness vs chat, same model, 400 instances
Pass@1 leader
41.2
Loop + Claude Code. Review ≠ work.
$ / resolved
$2.40–7.10
OpenCode cheapest. Droids dearest. Pass is not $.
Scaffolding premium
10–36 pt
Same weights, different harness. Literature, Aug 2026.
ACES Skill Lift
0.21
NVIDIA, 947 paired cases · 72.8% positive · ρ=0.14 vs scans
Review share of tokens
59.4%
Tokenomics arXiv:2601.14470 · ChatDev + GPT-5, n=30
Coord. at 4,280
15.4%
Hierarchical. Mesh would be a cancellation.
Unattributed overrun
35–50%
Zylos 2026, 5+ agents, no per-agent identity
Projects cancelled
40%
Gartner, by 2027 — cost, scale, risk
01 · HM-SWE-Pro 400
Same model. Ten shells. One score.
Chat without a harness is 18.2 pass@1. The best listed loop is 41.2. That gap is not a bigger context window. It is permissions, recovery, a spec, and a review agent that is not the author. Dollars per resolved instance do not rank the same as pass@1 — OpenCode wins the bill, Loop+Claude wins the work.
Pass@1
USD / resolved
| Harness | Pass@1 | $ / resolved | M tok | Recover | p50 min | |
|---|---|---|---|---|---|---|
| Loop + Claude Code | 41.2 | $4.80 | 1.12 | 91% | 18 | Review harness ≠ work harness. |
| Claude Code 2.4.1 | 38.6 | $5.40 | 1.31 | 88% | 16 | Managed default. Dear tokens, high recover. |
| Grok Build 0.9.2 | 37.1 | $4.10 | 1.18 | 86% | 14 | Worktree pool. Fast p50. |
| Codex CLI 1.8.4 | 35.4 | $3.90 | 1.24 | 84% | 15 | Open shell, commercial runtime. |
| Factory Droids 2.0 | 34.0 | $7.10 | 1.62 | 81% | 28 | Finishes tickets. Expensive hour. |
| Cursor Agent 1.6.2 | 32.8 | $6.20 | 1.55 | 79% | 22 | Wall-clock inflated by humans in the IDE. |
| OpenCode 0.22 | 29.4 | $2.40 | 1.48 | 74% | 19 | Cheapest resolved. Not the most resolved. |
| Gemini CLI 0.7.1 | 28.9 | $3.20 | 1.36 | 72% | 17 | Model-bound. Fine in a Gemini shop. |
| Copilot Workspace | 31.1 | $4.60 | 1.29 | 83% | 21 | SHA on every span. Forge-native. |
| Chat, no harness | 18.2 | $2.10 | 0.84 | 41% | 11 | The 2024 product. Not an agent. |
Directional finding cited: SWE-bench Pro analysis, Aug 2026. Per-row figures are the desk pack, labelled as such. See method.
02 · Cost
The cheap resolved instance is not the accepted one.
Scatter: pass@1 against dollars per resolved. The useful frontier is up and left. Chat sits low and cheap. Droids sit high and dear. Loop+Claude is the point you actually want if the unit of work is a merge, not a token.
Pass@1 vs $ / resolved
Recover vs pass@1
Recovery is failed tool calls that still produced a later patch. Chat sits low on both.
Session spend · delivery-core, 7d
The $0.75 cap sits after the bulk of the mass. The right tail is the cancellation. 1.6% of sessions above $8 is most of the surprise on the invoice.
Retry cascade
Each extra tool loop accumulates history. By six, half the sessions are already a worse deal than abort. The guard is a six-iteration default because the curve told us to.
Cost overrun vs. identity
Zylos 2026: organisations running five or more agents without per-agent attribution reported 35–50% overruns. The desk reproduces the shape. Mesh without a supervisor is how a project dies.
03 · Attribution
Token share is not contribution.
Steel dots are token %. Brass dots are accepted-work %. The line between them is the argument. Meridian Loop sits far to the right of its tokens. Cursor sits the other way — it holds the queue. A single pie would fire the wrong harness.
HMX · industry tokens, four quarters
Claude Code concentrates. Grok Build is the only lab still climbing in incoming telemetry. Internal harnesses are under-counted.
Meridian 7d · six meters
| Tok | Acc | Shapley | Liab. | |
|---|---|---|---|---|
| claude-code | 34.2% | 28.1% | 30.4% | 41.2% |
| codex-cli | 21.8% | 19.6% | 18.8% | 18.4% |
| grok-build | 16.4% | 18.2% | 17.6% | 14.1% |
| cursor-agent | 11.1% | 9.4% | 8.1% | 12.8% |
| opencode | 8.2% | 5.8% | 6.2% | 2.1% |
| meridian-loop | 8.3% | 18.9% | 18.9% | 11.4% |
04 · Scale
One thousand is operations. Ten thousand is a market.
Coordination overhead against agent count. Mesh (alert) is a toy above 80. Hub-spoke (steel) is the default to ~800, then the hub is the incident. Hierarchy (brass) is what Meridian runs at 4,280. Pipeline is a happy path that does not survive a real ticket mix.
| n | pipeline ovh | hub-spoke ovh | hierarchical ovh | mesh ovh | Hier. $/h |
|---|---|---|---|---|---|
| 80 | 5.1% | 8.4% | 6.2% | 51.5% | $4 |
| 800 | 6.4% | 11.6% | 7.8% | 72.0% | $41 |
| 1,000 | 6.8% | 12.5% | 8.2% | 72.0% | $52 |
| 4,280 | 12.7% | 27.3% | 15.4% | 72.0% | $237 |
| 10,000 | 22.0% | 42.0% | 28.0% | 72.0% | $614 |
Projector, not a measurement of your fleet. Open it on the desk and drag n to 10,000.
Scale projector05 · Skills & advances
Most skills do not move pass@1. A few move the bill.
Ablation on Claude Code, same model, same 400. Plan-work-review is the spec. The KV-cache scheduler does not add a point of pass and deletes 18% of dollars. The security gate does not show up on SWE and is still not optional. Eval-pack-swe is a cost you pay to know any of this.
Pass@1 through the stack
Accepted-work lift (production, 7d)
| Skill / advance | Accepted lift | Pass pt | Cost Δ | Invocations |
|---|---|---|---|---|
| plan-work-review | 12.8% | +4.1 | 3% | 18,420 |
| repo-memory | 4.6% | +1.4 | 1% | 6,200 |
| retry-cascade-guard | 2.1% | +1.2 | -16% | 940 |
| security-review-gate | 3.2% | 0.0 | 5% | 18,110 |
| kv-cache-scheduler | 0.4% | 0.0 | -18% | 4,100 |
| cost-circuit | 0.1% | -0.3 | -22% | 940 |
| eval-pack-swe | 0.0% | 0.0 | 8% | 860 |
| progressive-disclosure | 1.8% | +0.6 | -9% | 140,000 |
06 · Ecosystem
Agents are the small number. Skills and MCP are the market.
June 2026: skills.sh GA, 600k OSS skills (Vercel). August 2026: Skillful.sh 55-directory aggregate, 553,855 tools (skills 58.2% / MCP 37.5% / agents 4.3%). The June skills.sh figure is a single registry; Skillful is a union. Do not subtract them.
Index, Dec 2025 = 100. June skills.sh 600k is a different universe from the August Skillful union — the dip is not a crash.
| Month | Skills | MCP | Agents | Tools (union) |
|---|---|---|---|---|
| 2025-12 | 84,000 | 92,000 | 9,100 | 185,000 |
| 2026-02 | 210,000 | 128,000 | 12,400 | 350,000 |
| 2026-04 | 268,000 | 164,000 | 16,800 | 449,000 |
| 2026-06 | 600,000 | 188,000 | 20,100 | 808,000 |
| 2026-08 | 322,131 | 207,698 | 24,026 | 553,855 |
08 · Skills, MCP, context
A skill is not a scan. MCP is not solved.
NVIDIA ACES (arXiv:2608.20614) is the first large paired with/without skill protocol: mean composite Skill Lift 0.2134 on 947 cases, positive in 72.8%. Scan-only gates correlate with live judges at ρ=0.14. Jiang et al. show why: skills mostly stabilize action (procedural anchoring 65.7%), and retrieval dies as the pool grows (29.6% precision at 5 skills, 3.3% at 100). MCP-Universe: frontier models still fail more than half the time on real servers.
ACES · Skill Lift by dimension
NVIDIA blog snapshot, Aug 2026 · +0.21 composite · CI 0.1967–0.2301
Jiang et al. · retrieval precision vs pool
arXiv:2608.14036 · anchoring 65.7% of cases · knowledge 4.5% · vs Workflow Memory +6.06 pt
MCP-Universe success %
MCP-Universe: 6 domains, 11 real MCP servers, >20 models. Execution-based evaluators. Cursor and Claude Code did not beat ReAct on this board. Long-context and unknown-tool failure modes dominate.
09 · Token tax
The bill is review, input, and coordination — not generation.
Salim et al., arXiv:2601.14470. ChatDev + GPT-5, 30 SDLC tasks. Code Review 59.4% ± 7.9% of tokens; initial Coding 8.6% ± 1.4%. Input 53.9% overall. The bill is refinement, not generation. Sun et al. (arXiv:2608.22152, EMNLP 2026) name the rest a collaboration tax: structured, capability-monotonic, partly tractable. PACT, file-mediated coordination, LATTE, and Shamay’s production patterns are the interventions. Hybrid multi-agent topologies spend 5× the tokens of a single agent for no extra success.
Where tokens go · Tokenomics n=30
Coordination interventions
MAS turn overhead vs SAS
Nature Machine Intelligence / MIT (2026). Reasoning-turn overhead vs a single-agent system under matched compute ceilings. Hybrid uses 6.2× the turns of SAS for statistically indistinguishable success versus centralized. Successes per 1,000 tokens: SAS 67.7, hybrid 13.6 (5.0× worse).
| Architecture | Turn overhead | Successes / 1k tok |
|---|---|---|
| Single-agent system | — | 67.7 |
| Independent MAS | +58% | — |
| Decentralized | +263% | 23.9 |
| Centralized | +285% | 21.5 |
| Hybrid | +515% | 13.6 |
10 · Live boards · 29 Aug 2026
Verified is saturating. Pro and Terminal-Bench 3.0 are not.
Verified (BenchLM, 29 Aug 2026) is saturating — top three within 1.0 pt. Scale live Pro public (mini-swe-agent, uncapped, 250 turns) is a different object from the ~23% paper-era Pro snapshots. Terminal-Bench 2.0 is also saturating; 3.0 (74 tasks, 7 domains) is not. TB 2.1 (Artificial Analysis, Terminus 2) zero-scores reward hacks. Do not mix boards. SWE-bench Pro is 1,865 tasks, 41 repos (731 public / 858 held-out / 276 commercial). GPT-5.6 Sol’s 91.9% on TB 2.0 is 88.8% as a solo agent and 91.9% when “ultra” coordinates four agents — a harness number printed as a model number.
SWE-Verified · BenchLM
SWE-Pro public live · Scale mini-swe-agent
Terminal-Bench 2.0
Terminal-Bench 2.1 · Terminus 2
Artificial Analysis. Reward-hack zero. Do not mix with 2.0 or 3.0.
Terminal-Bench 3.0 · 74 tasks
11 · Working mix
Do not pick a winner. Pick a mix you can attribute.
The eval does not say “buy Claude Code and go home.” It says: concentrate workers on the harnesses that resolve, keep a promotion harness that will look small on tokens, keep OpenCode for the jobs that should be cheap, and never let the author review. Meridian’s live mix is the existence proof. The board below is the slightly tightened version the numbers support for a 4k delivery fleet in this window.
Suggested token mix · 4k delivery
- claude-code32.0%
- codex-cli20.0%
- grok-build18.0%
- meridian-loop10.0%
- opencode12.0%
- cursor-agent8.0%
Workers are a mix, not a brand. Claude Code for the hard tickets, Codex and Grok for volume, OpenCode for eval and the jobs that should be cheap.
Promotion (Loop) is small on tokens and large on accepted work. If your token pie matches your accepted pie, you do not have a promotion harness.
Cursor stays in the IDE band. It is not a 4,000-agent worker.
Review is a different harness from work. That constraint is worth more pass@1 than a model upgrade in this window.
Above 800 workers: hierarchical, Helix, $0.75 session cap, prodOnlySigned. The eval does not save you from topology.
12 · Landscape pack
Same method. More shells. Still desk.
The original ten stay the headline. These nine are the rest of the floor that operators actually install — Amp, Kiro, OpenHands, Aider, Junie, Cline, Goose, Qwen Code, Crush. Same window, same model hold, same 400. Conservative. Labelled desk. Not Scale live Pro, not Verified.
Pass@1 · landscape shells
USD / resolved
| Harness | Pass@1 | $ / resolved | Recover | p50 min | |
|---|---|---|---|---|---|
| Amp | 36.4 | $5.80 | 85% | 19 | Repo intel in the loop. Dear tokens, high recover. |
| Kiro | 34.8 | $4.40 | 82% | 23 | Spec-first. Property tests lift recover, not p50. |
| OpenHands | 33.6 | $3.40 | 80% | 20 | Docker sandbox. Cited Verified is higher; Pro is harder. |
| Aider | 31.4 | $2.80 | 76% | 17 | Git-native. Cheap resolved. Recovers less than the labs. |
| Junie | 30.8 | $4.00 | 77% | 24 | IDE-bound. Wall-clock inflated like Cursor. |
| Cline | 30.2 | $3.60 | 73% | 21 | Plugin marketplace. Loop quality follows the plugin. |
| Goose | 29.8 | $2.90 | 71% | 16 | Local-first. Recipes help; no promotion harness. |
| Qwen Code | 27.6 | $2.20 | 68% | 15 | Open weights in the shell. Cheap. Not the top of Pro. |
| Crush | 26.4 | $2.60 | 64% | 13 | Charm's TUI. Fast p50, thin recover. |
Recover vs pass · landscape
13 · Method & sources
What we will claim, and what we will not.
HM-SWE-Pro 400 is a desk subset of SWE-bench Pro. Model held constant (frontier snapshot 2026-08-18). Harness is the independent variable. Three seeds. Pass@1.
The directional finding — harness swap moved pass@1 more than many model upgrades — is cited to the August 2026 SWE-bench Pro analysis (AINews, 8 Aug 2026, via @joelniklaus). Per-harness figures in the table are ours, labelled as such, and will move when the pack is re-run.
Cost is USD at posted lab list for that window, excluding seat licenses, including hosted agent-hours where the listing is usage-priced. Recovery is share of failed tool calls that produced a subsequent successful patch in-session.
Meridian mix and skill lift are trailing-7-day production, not the eval subset. Shapley-lite is 2% weekly ablation on the ticket corpus. Token share is not contribution.
Scale curves are the desk projector (see /desk/fleet): mesh overhead from channel count, hub-spoke and hierarchy from n. They are a model, not a measurement of your fleet.
Ecosystem counts: Skillful.sh 29 Aug 2026 (55 directories); skills.sh GA 5 Jun 2026; Agent Skills open directory May 2026 (1.2M+). Different universes. Cited, not blended.
External scaffolding-premium rows are named-source measurements (Scale SEAL, Augment, Endor Labs, Stanford IRIS, LangChain, Lewis arXiv:2608.26218). They do not share a harness, a split, or a week. The range 10–36 pt is the phenomenon. Any one row is a citation, not a ranking.
Leaderboards (SWE-Verified, Scale live Pro, Terminal-Bench 2.0 / 2.1 / 3.0) are different objects. Verified is saturating. Paper-era Pro ~23% is not the live mini-swe-agent board. We print both and refuse to subtract them.
The landscape extension (Amp, Kiro, OpenHands, Aider, Junie, Cline, Goose, Qwen Code, Crush) uses the same desk method and window as HM-SWE-Pro 400. Figures are conservative and labelled desk. They are not a live Scale board.
Open-index mix (kind, via, subscribe, trust) is counted from the listing catalog at render. Editorial install counts stay editorial. Unhoused slips (community bucket, generic MCP) are not assigned a fake studio.
ACES Skill Lift (arXiv:2608.20614) is paired with/without on a fixed task, harness, workspace, scorer. It is the closest published cousin of this desk’s skill ablation. Scan-only skill gates (Spearman ρ=0.14 vs live judges) are not a substitute.
| ID | Who | Year | What we took |
|---|---|---|---|
| 2608.26218 | Lewis | 2026 | Same Model, Different Harness. Tight-window Verified: F2PF 28→49, complete 43→72. |
| 2608.20614 | Kevin et al. (NVIDIA ACES) | 2026 | Skill Lift 0.2134 on 947 paired cases. Scan-only vs judge ρ=0.14. |
| 2608.14036 | Jiang et al. (UCSD) | 2026 | Skills as procedural anchors (65.7%). Retrieval precision 29.6%→3.3% as pool 5→100. |
| 2608.22152 | Sun et al. | 2026 | Collaboration tax. 32 tasks, 11 models, 7 providers. EMNLP 2026. |
| 2608.16801 | Aste | 2026 | File-mediated coordination −42% output tokens at 8 agents on message-heavy work. |
| 2608.17188 | Shamay | 2026 | Six production token patterns. 60–70% cut. 2,420 confirmatory trials. |
| 2601.14470 | Salim et al. | 2026 | Tokenomics. ChatDev+GPT-5, 30 tasks. Review 59.4% of tokens. Input 53.9%. |
| 2606.17819 | Skill-eval @ KDD workshop | 2026 | 500 skills, 1,000 tasks, 19 model configs. Instruction-following varies widely. |
| 2606.11435 | Ding et al. | 2026 | Survey: skill evaluation and evolution. Six benchmark categories. |
| 2602.12430 | Skills survey | 2026 | Architecture, acquisition, security. 26.1% of public skills flagged in empirical analyses. |
| 2512.24565 | Liu et al. MCPAgentBench | 2026 | Real-world MCP tool-use with distractors in the candidate list. |
| 2606.05304 | PACT | 2026 | Action-state communication. −38.7% tokens; SWE-agent input −50.4%. |
| 2603.15183 | Parakhin | 2026 | Token Coherence / MESI for artifacts. Simulated 84–95% sync savings. |
| 2605.06320 | LATTE | 2026 | Adaptive task graphs. Token cost 47.5% of next-best. Overwrites 4.3 vs 22.8–35.4. |
| 2606.17454 | Gupta et al. SSA | 2026 | Simple Strands Agent. Minimal harness, SOTA-level Verified / Pro / TB2. |
| SEAL-Pro | Scale SEAL | 2026 | SWE-bench Pro public, SWE-agent / mini-swe-agent. Live board ≠ paper-era ~23%. |
| IRIS-MH | Stanford IRIS Meta-Harness | 2026 | Opus 4.6 on TB2.0: 43% → 76.4% by harness only. |
| MCP-U | MCP-Universe | 2026 | GPT-5-High 44.16% / Grok-4 33.33% on 11 live MCP servers. |
| NMI-MAS | Nature Mach. Intell. / MIT | 2026 | MAS turn overhead 58–515%. SAS 67.7 successes/1k tok vs hybrid 13.6. |
| Skillful | Skillful.sh | 2026-08-29 | 553,855 tools across 55 directories. Skills 58.2 / MCP 37.5 / agents 4.3. |