HarnessmarketEnter the desk
/HMX Claude Code 29.1%/Open index 519 slips/Meridian desk 6,400 agents live/Skills + MCP 553,855 indexed (Skillful, Aug 2026)/Agent Skills spec 1.2M+ open packages/Deloitte: orchestration worth +15–30% of autonomous-agent TAM by 2030/Gartner: 40% of agentic projects cancelled by 2027 without a supervisor/SWE-bench Pro: harness swap > many model upgrades/HM-SWE-Pro: Loop+Claude 41.2 pass@1 · chat 18.2 · same model/Lewis 2608.26218: F2PF 28→49 under a tighter harness, same model/ACES: Skill Lift 0.21 · 947 paired cases · scan vs live ρ=0.14/Tokenomics: code review 59.4% of ChatDev tokens/HMX Claude Code 29.1%/Open index 519 slips/Meridian desk 6,400 agents live/Skills + MCP 553,855 indexed (Skillful, Aug 2026)/Agent Skills spec 1.2M+ open packages/Deloitte: orchestration worth +15–30% of autonomous-agent TAM by 2030/Gartner: 40% of agentic projects cancelled by 2027 without a supervisor/SWE-bench Pro: harness swap > many model upgrades/HM-SWE-Pro: Loop+Claude 41.2 pass@1 · chat 18.2 · same model/Lewis 2608.26218: F2PF 28→49 under a tighter harness, same model/ACES: Skill Lift 0.21 · 947 paired cases · scan vs live ρ=0.14/Tokenomics: code review 59.4% of ChatDev tokens

Eval desk · through 2026-08-29

Hold the model. The harness is the independent variable.

On SWE-bench Pro, swapping the harness moved pass@1 more than many model upgrades. The literature now says the same thing in a dozen independent runs: 10–36 points of harness, Skill Lift near 0.21, review eating 59% of tokens, MCP success still under 45% at the frontier. Desk pack plus sources — labelled, not blended.

HM-SWE-Pro 400arXiv 2608.26218ACES 2608.20614Tokenomics 2601.14470Thesis brief

HM-SWE-Pro Δ

+22.4 pt

Best harness vs chat, same model, 400 instances

Pass@1 leader

41.2

Loop + Claude Code. Review ≠ work.

$ / resolved

$2.40–7.10

OpenCode cheapest. Droids dearest. Pass is not $.

Scaffolding premium

10–36 pt

Same weights, different harness. Literature, Aug 2026.

ACES Skill Lift

0.21

NVIDIA, 947 paired cases · 72.8% positive · ρ=0.14 vs scans

Review share of tokens

59.4%

Tokenomics arXiv:2601.14470 · ChatDev + GPT-5, n=30

Coord. at 4,280

15.4%

Hierarchical. Mesh would be a cancellation.

Unattributed overrun

35–50%

Zylos 2026, 5+ agents, no per-agent identity

Projects cancelled

40%

Gartner, by 2027 — cost, scale, risk

01 · HM-SWE-Pro 400

Same model. Ten shells. One score.

Chat without a harness is 18.2 pass@1. The best listed loop is 41.2. That gap is not a bigger context window. It is permissions, recovery, a spec, and a review agent that is not the author. Dollars per resolved instance do not rank the same as pass@1 — OpenCode wins the bill, Loop+Claude wins the work.

Pass@1

Loop + Claude Code
41.2%
Claude Code 2.4.1
38.6%
Grok Build 0.9.2
37.1%
Codex CLI 1.8.4
35.4%
Factory Droids 2.0
34.0%
Cursor Agent 1.6.2
32.8%
OpenCode 0.22
29.4%
Gemini CLI 0.7.1
28.9%
Copilot Workspace
31.1%
Chat, no harness
18.2%

USD / resolved

Chat, no harness
2.10
OpenCode 0.22
2.40
Gemini CLI 0.7.1
3.20
Codex CLI 1.8.4
3.90
Grok Build 0.9.2
4.10
Copilot Workspace
4.60
Loop + Claude Code
4.80
Claude Code 2.4.1
5.40
Cursor Agent 1.6.2
6.20
Factory Droids 2.0
7.10
HarnessPass@1$ / resolvedM tokRecoverp50 min
Loop + Claude Code41.2$4.801.1291%18Review harness ≠ work harness.
Claude Code 2.4.138.6$5.401.3188%16Managed default. Dear tokens, high recover.
Grok Build 0.9.237.1$4.101.1886%14Worktree pool. Fast p50.
Codex CLI 1.8.435.4$3.901.2484%15Open shell, commercial runtime.
Factory Droids 2.034.0$7.101.6281%28Finishes tickets. Expensive hour.
Cursor Agent 1.6.232.8$6.201.5579%22Wall-clock inflated by humans in the IDE.
OpenCode 0.2229.4$2.401.4874%19Cheapest resolved. Not the most resolved.
Gemini CLI 0.7.128.9$3.201.3672%17Model-bound. Fine in a Gemini shop.
Copilot Workspace31.1$4.601.2983%21SHA on every span. Forge-native.
Chat, no harness18.2$2.100.8441%11The 2024 product. Not an agent.

Directional finding cited: SWE-bench Pro analysis, Aug 2026. Per-row figures are the desk pack, labelled as such. See method.

02 · Cost

The cheap resolved instance is not the accepted one.

Scatter: pass@1 against dollars per resolved. The useful frontier is up and left. Chat sits low and cheap. Droids sit high and dear. Loop+Claude is the point you actually want if the unit of work is a merge, not a token.

Pass@1 vs $ / resolved

Loop+CCClaudeGrokCodexFactoryCursorOpenCodeGeminiCopilotChat,USD per resolved instancePass@1

Recover vs pass@1

Recovery is failed tool calls that still produced a later patch. Chat sits low on both.

Loop+CCClaudeGrokCodexFactoryCursorOpenCodeGeminiCopilotChat,Recover %Pass@1

Session spend · delivery-core, 7d

The $0.75 cap sits after the bulk of the mass. The right tail is the cancellation. 1.6% of sessions above $8 is most of the surprise on the invoice.

<$0.10
41.2%
$0.10–0.30
27.1%
$0.30–0.75
19.3%
$0.75–2
7.1%
$2–8
3.7%
>$8
1.6%

Retry cascade

Each extra tool loop accumulates history. By six, half the sessions are already a worse deal than abort. The guard is a six-iteration default because the curve told us to.

1 loops
0.11
2 loops
0.24
3 loops
0.41
4 loops
0.68
5 loops
1.12
6 loops
1.84
8 loops
4.60
12 loops
14.20

Cost overrun vs. identity

Zylos 2026: organisations running five or more agents without per-agent attribution reported 35–50% overruns. The desk reproduces the shape. Mesh without a supervisor is how a project dies.

1 harness, attributed
4%
2–4, attributed
11%
5+, attributed
18%
5+, no identity on spans
42%
mesh >80, no supervisor
67%

03 · Attribution

Token share is not contribution.

Steel dots are token %. Brass dots are accepted-work %. The line between them is the argument. Meridian Loop sits far to the right of its tokens. Cursor sits the other way — it holds the queue. A single pie would fire the wrong harness.

Tokens Accepted
claude-code34.228.1codex-cli21.819.6grok-build16.418.2cursor-agent11.19.4opencode8.25.8meridian-loop8.318.9

HMX · industry tokens, four quarters

02550751002025 Q42026 Q3
claude-codecodex-clicursor-agentcopilot-workspacegrok-buildgemini-cliopencodemeridian-loop

Claude Code concentrates. Grok Build is the only lab still climbing in incoming telemetry. Internal harnesses are under-counted.

Meridian 7d · six meters

TokAccShapleyLiab.
claude-code34.2%28.1%30.4%41.2%
codex-cli21.8%19.6%18.8%18.4%
grok-build16.4%18.2%17.6%14.1%
cursor-agent11.1%9.4%8.1%12.8%
opencode8.2%5.8%6.2%2.1%
meridian-loop8.3%18.9%18.9%11.4%
Live attribution

04 · Scale

One thousand is operations. Ten thousand is a market.

Coordination overhead against agent count. Mesh (alert) is a toy above 80. Hub-spoke (steel) is the default to ~800, then the hub is the incident. Hierarchy (brass) is what Meridian runs at 4,280. Pipeline is a happy path that does not survive a real ticket mix.

pipelinehub-spokehierarchicalmesh
0.0%19%38%57%76%10 agents10,000
npipeline ovhhub-spoke ovhhierarchical ovhmesh ovhHier. $/h
805.1%8.4%6.2%51.5%$4
8006.4%11.6%7.8%72.0%$41
1,0006.8%12.5%8.2%72.0%$52
4,28012.7%27.3%15.4%72.0%$237
10,00022.0%42.0%28.0%72.0%$614

Projector, not a measurement of your fleet. Open it on the desk and drag n to 10,000.

Scale projector

05 · Skills & advances

Most skills do not move pass@1. A few move the bill.

Ablation on Claude Code, same model, same 400. Plan-work-review is the spec. The KV-cache scheduler does not add a point of pass and deletes 18% of dollars. The security gate does not show up on SWE and is still not optional. Eval-pack-swe is a cost you pay to know any of this.

Pass@1 through the stack

Claude Code only
38.6
+ plan-work-review
42.7
+ Meridian Loop (review ≠ work)
44.9
+ KV-cache scheduler
44.9
+ retry-cascade guard
46.1
+ cost-circuit $0.75
45.8
+ security-review-gate
45.8

Accepted-work lift (production, 7d)

plan-work-review
12.8%
repo-memory
4.6%
retry-cascade-guard
2.1%
security-review-gate
3.2%
kv-cache-scheduler
0.4%
cost-circuit
0.1%
eval-pack-swe
0.0%
progressive-disclosure
1.8%
Skill / advanceAccepted liftPass ptCost ΔInvocations
plan-work-review12.8%+4.13%18,420
repo-memory4.6%+1.41%6,200
retry-cascade-guard2.1%+1.2-16%940
security-review-gate3.2%0.05%18,110
kv-cache-scheduler0.4%0.0-18%4,100
cost-circuit0.1%-0.3-22%940
eval-pack-swe0.0%0.08%860
progressive-disclosure1.8%+0.6-9%140,000

06 · Ecosystem

Agents are the small number. Skills and MCP are the market.

June 2026: skills.sh GA, 600k OSS skills (Vercel). August 2026: Skillful.sh 55-directory aggregate, 553,855 tools (skills 58.2% / MCP 37.5% / agents 4.3%). The June skills.sh figure is a single registry; Skillful is a union. Do not subtract them.

0.018837556375025-1226-0226-0426-0626-08
skillsmcpagents

Index, Dec 2025 = 100. June skills.sh 600k is a different universe from the August Skillful union — the dip is not a crash.

MonthSkillsMCPAgentsTools (union)
2025-1284,00092,0009,100185,000
2026-02210,000128,00012,400350,000
2026-04268,000164,00016,800449,000
2026-06600,000188,00020,100808,000
2026-08322,131207,69824,026553,855

07 · Literature

The scaffolding premium is a measured object.

Same weights, different shell. Scale’s standardized SWE-agent run of Opus 4.5 sits at 45.9% on Pro public. Three vendor harnesses on the same model land 49.8–51.8. Endor Labs moved GPT-5.5 25.7 points by swapping Codex’s native loop for Cursor. Stanford IRIS moved Opus 4.6 33.4 points on Terminal-Bench 2.0. The trade press’s 10–20 point “routine gap” is the conservative read.

Δ pass, harness only

Minimal → Claude Code
36.0 pt
Stanford IRIS Meta-Harness
33.4 pt
Claw-SWE-Bench max gap
27.4 pt
Codex native → Cursor
25.7 pt
Qwen3-32B scaffolds
20.0 pt
LangChain harness swap
13.7 pt
Auggie
5.9 pt
Cursor
4.3 pt
Claude Code
3.9 pt

Lo → hi, same weights

lo / baselinehi / treatment
Scale SEAL / SWE-agent45.945.9Claude Code45.949.8Cursor45.950.2Auggie45.951.8Codex native → Cursor61.587.2Minimal → Claude Code42.078.0LangChain harness swap52.866.5Claw-SWE-Bench max gap0.027.4Stanford IRIS Meta-Harness43.076.4Qwen3-32B scaffolds28.048.0
ComparisonModelBenchLoHiΔSource
Scale SEAL / SWE-agentClaude Opus 4.5SWE-bench Pro public45.945.9Scale SEAL; FutureAGI 2026-07-04
Claude CodeClaude Opus 4.5SWE-bench Pro public45.949.8+3.9Augment Code table via FutureAGI
CursorClaude Opus 4.5SWE-bench Pro public45.950.2+4.3Augment Code table via FutureAGI
AuggieClaude Opus 4.5SWE-bench Pro public45.951.8+5.9Augment Code table via FutureAGI
Codex native → CursorGPT-5.5SWE-bench61.587.2+25.7Endor Labs; LeCompute 2026-07-13
Minimal → Claude CodeClaude OpusIndependent coding4278+36LeCompute 2026-07-13
LangChain harness swapheld constantTerminal-Bench52.866.5+13.7LangChain; LeCompute 2026-07-13
Claw-SWE-Bench max gapheld constant350 tasks27.4+27.4Claw-SWE-Bench via LeCompute
Stanford IRIS Meta-HarnessClaude Opus 4.6Terminal-Bench 2.04376.4+33.4Stanford IRIS; Genαi 2026-07-11
Qwen3-32B scaffoldsQwen3-32BSWE-bench (Agentless / OpenHands / SWE-Agent)2848+20AgentMarketCap 2026-04-07

Lewis 2026 · arXiv:2608.26218

Same Model, Different Harness: Different Coding-Agent Results

Control: full conversation in time order. Treatment: mechanically shorten older tool results as context fills; respond to stall. Frozen treatment also lifts three other models with no retuning. Wide-window Qwen3.6 arms are close on Verified and Pro.

F2PF

2849

Complete

4372

169 Verified tasks · 20,480 tok window · 480s cap · no model retune

Agent Lightning · 2026-08-18

RL inside the harness, not instead of it

RL on ~6,000 examples inside the existing harness loop. The gain is model+harness+data+eval, not a library credit.

41.856.4+14.6 pt

Qwen3.5-9B · SWE-bench Verified

08 · Skills, MCP, context

A skill is not a scan. MCP is not solved.

NVIDIA ACES (arXiv:2608.20614) is the first large paired with/without skill protocol: mean composite Skill Lift 0.2134 on 947 cases, positive in 72.8%. Scan-only gates correlate with live judges at ρ=0.14. Jiang et al. show why: skills mostly stabilize action (procedural anchoring 65.7%), and retrieval dies as the pool grows (29.6% precision at 5 skills, 3.3% at 100). MCP-Universe: frontier models still fail more than half the time on real servers.

ACES · Skill Lift by dimension

NVIDIA blog snapshot, Aug 2026 · +0.21 composite · CI 0.19670.2301

Correctness
41 pt
Discoverability
40 pt
Effectiveness
39 pt
Efficiency
35 pt
Security
1 pt

Jiang et al. · retrieval precision vs pool

arXiv:2608.14036 · anchoring 65.7% of cases · knowledge 4.5% · vs Workflow Memory +6.06 pt

pool 5
29.6%
pool 20
14.1%
pool 50
7.2%
pool 100
3.3%

MCP-Universe success %

MCP-Universe: 6 domains, 11 real MCP servers, >20 models. Execution-based evaluators. Cursor and Claude Code did not beat ReAct on this board. Long-context and unknown-tool failure modes dominate.

GPT-5-High
44.2%
Grok-4
33.3%
Claude-4.0-Sonnet
29.4%

09 · Token tax

The bill is review, input, and coordination — not generation.

Salim et al., arXiv:2601.14470. ChatDev + GPT-5, 30 SDLC tasks. Code Review 59.4% ± 7.9% of tokens; initial Coding 8.6% ± 1.4%. Input 53.9% overall. The bill is refinement, not generation. Sun et al. (arXiv:2608.22152, EMNLP 2026) name the rest a collaboration tax: structured, capability-monotonic, partly tractable. PACT, file-mediated coordination, LATTE, and Shamay’s production patterns are the interventions. Hybrid multi-agent topologies spend 5× the tokens of a single agent for no extra success.

Where tokens go · Tokenomics n=30

Code Review59.458.9%
Coding8.68.5%
Design2.42.4%
Testing (when run)10.310.2%
Documentation20.119.9%
Input53.954.0%
Output24.424.4%
Reasoning21.621.6%

Coordination interventions

3-agent vs solo pipeline
190.0%
PACT mean token cut
-38.7%
PACT on SWE-agent input
-50.4%
File-mediated @ 8 agents
-42.0%
Forced files on chained tasks
13.5%
LATTE vs next-best tokens
-52.5%
Shamay production token cut
-65.0%

MAS turn overhead vs SAS

Nature Machine Intelligence / MIT (2026). Reasoning-turn overhead vs a single-agent system under matched compute ceilings. Hybrid uses 6.2× the turns of SAS for statistically indistinguishable success versus centralized. Successes per 1,000 tokens: SAS 67.7, hybrid 13.6 (5.0× worse).

Single-agent system
0%
Independent MAS
58%
Decentralized
263%
Centralized
285%
Hybrid
515%
ArchitectureTurn overheadSuccesses / 1k tok
Single-agent system67.7
Independent MAS+58%
Decentralized+263%23.9
Centralized+285%21.5
Hybrid+515%13.6

10 · Live boards · 29 Aug 2026

Verified is saturating. Pro and Terminal-Bench 3.0 are not.

Verified (BenchLM, 29 Aug 2026) is saturating — top three within 1.0 pt. Scale live Pro public (mini-swe-agent, uncapped, 250 turns) is a different object from the ~23% paper-era Pro snapshots. Terminal-Bench 2.0 is also saturating; 3.0 (74 tasks, 7 domains) is not. TB 2.1 (Artificial Analysis, Terminus 2) zero-scores reward hacks. Do not mix boards. SWE-bench Pro is 1,865 tasks, 41 repos (731 public / 858 held-out / 276 commercial). GPT-5.6 Sol’s 91.9% on TB 2.0 is 88.8% as a solo agent and 91.9% when “ultra” coordinates four agents — a harness number printed as a model number.

SWE-Verified · BenchLM

Claude Opus 5
96.0%
Claude Mythos 5
95.5%
Claude Fable 5
95.0%
Claude Opus 4.8
88.6%
GPT-5.3 Codex
85.0%

SWE-Pro public live · Scale mini-swe-agent

Muse Spark 1.1*
61.5%
gpt-5.4 (xHigh)*
59.1%
Muse Spark*
55.0%
claude-opus-4-6 (thinking)*
51.9%

Terminal-Bench 2.0

GPT-5.6 Sol
91.9%
Claude Mythos 5
88.0%
GPT-5.6 Terra
87.4%
Grok 4.5
83.3%
GPT-5.5
82.0%

Terminal-Bench 2.1 · Terminus 2

Artificial Analysis. Reward-hack zero. Do not mix with 2.0 or 3.0.

GPT-5.6 Sol (max)
89.5%
Claude Opus 5 (max)
89.1%
Grok 4.6 (high)
88.4%

Terminal-Bench 3.0 · 74 tasks

Claude Opus 5
42.7%
GPT-5.6 Sol
34.6%
Claude Fable 5
34.0%
Grok 4.5
15.7%

11 · Working mix

Do not pick a winner. Pick a mix you can attribute.

The eval does not say “buy Claude Code and go home.” It says: concentrate workers on the harnesses that resolve, keep a promotion harness that will look small on tokens, keep OpenCode for the jobs that should be cheap, and never let the author review. Meridian’s live mix is the existence proof. The board below is the slightly tightened version the numbers support for a 4k delivery fleet in this window.

Suggested token mix · 4k delivery

  • claude-code32.0%
  • codex-cli20.0%
  • grok-build18.0%
  • meridian-loop10.0%
  • opencode12.0%
  • cursor-agent8.0%
claude-code
32.0%
codex-cli
20.0%
grok-build
18.0%
meridian-loop
10.0%
opencode
12.0%
cursor-agent
8.0%

Workers are a mix, not a brand. Claude Code for the hard tickets, Codex and Grok for volume, OpenCode for eval and the jobs that should be cheap.

Promotion (Loop) is small on tokens and large on accepted work. If your token pie matches your accepted pie, you do not have a promotion harness.

Cursor stays in the IDE band. It is not a 4,000-agent worker.

Review is a different harness from work. That constraint is worth more pass@1 than a model upgrade in this window.

Above 800 workers: hierarchical, Helix, $0.75 session cap, prodOnlySigned. The eval does not save you from topology.

12 · Landscape pack

Same method. More shells. Still desk.

The original ten stay the headline. These nine are the rest of the floor that operators actually install — Amp, Kiro, OpenHands, Aider, Junie, Cline, Goose, Qwen Code, Crush. Same window, same model hold, same 400. Conservative. Labelled desk. Not Scale live Pro, not Verified.

Pass@1 · landscape shells

Amp
36.4%
Kiro
34.8%
OpenHands
33.6%
Aider
31.4%
Junie
30.8%
Cline
30.2%
Goose
29.8%
Qwen Code
27.6%
Crush
26.4%

USD / resolved

Qwen Code
2.20
Crush
2.60
Aider
2.80
Goose
2.90
OpenHands
3.40
Cline
3.60
Junie
4.00
Kiro
4.40
Amp
5.80
HarnessPass@1$ / resolvedRecoverp50 min
Amp36.4$5.8085%19Repo intel in the loop. Dear tokens, high recover.
Kiro34.8$4.4082%23Spec-first. Property tests lift recover, not p50.
OpenHands33.6$3.4080%20Docker sandbox. Cited Verified is higher; Pro is harder.
Aider31.4$2.8076%17Git-native. Cheap resolved. Recovers less than the labs.
Junie30.8$4.0077%24IDE-bound. Wall-clock inflated like Cursor.
Cline30.2$3.6073%21Plugin marketplace. Loop quality follows the plugin.
Goose29.8$2.9071%16Local-first. Recipes help; no promotion harness.
Qwen Code27.6$2.2068%15Open weights in the shell. Cheap. Not the top of Pro.
Crush26.4$2.6064%13Charm's TUI. Fast p50, thin recover.

Recover vs pass · landscape

AmpKiroOpenHandsAiderJunieClineGooseQwen CodeCrushRecover %Pass@1

13 · Method & sources

What we will claim, and what we will not.

HM-SWE-Pro 400 is a desk subset of SWE-bench Pro. Model held constant (frontier snapshot 2026-08-18). Harness is the independent variable. Three seeds. Pass@1.

The directional finding — harness swap moved pass@1 more than many model upgrades — is cited to the August 2026 SWE-bench Pro analysis (AINews, 8 Aug 2026, via @joelniklaus). Per-harness figures in the table are ours, labelled as such, and will move when the pack is re-run.

Cost is USD at posted lab list for that window, excluding seat licenses, including hosted agent-hours where the listing is usage-priced. Recovery is share of failed tool calls that produced a subsequent successful patch in-session.

Meridian mix and skill lift are trailing-7-day production, not the eval subset. Shapley-lite is 2% weekly ablation on the ticket corpus. Token share is not contribution.

Scale curves are the desk projector (see /desk/fleet): mesh overhead from channel count, hub-spoke and hierarchy from n. They are a model, not a measurement of your fleet.

Ecosystem counts: Skillful.sh 29 Aug 2026 (55 directories); skills.sh GA 5 Jun 2026; Agent Skills open directory May 2026 (1.2M+). Different universes. Cited, not blended.

External scaffolding-premium rows are named-source measurements (Scale SEAL, Augment, Endor Labs, Stanford IRIS, LangChain, Lewis arXiv:2608.26218). They do not share a harness, a split, or a week. The range 10–36 pt is the phenomenon. Any one row is a citation, not a ranking.

Leaderboards (SWE-Verified, Scale live Pro, Terminal-Bench 2.0 / 2.1 / 3.0) are different objects. Verified is saturating. Paper-era Pro ~23% is not the live mini-swe-agent board. We print both and refuse to subtract them.

The landscape extension (Amp, Kiro, OpenHands, Aider, Junie, Cline, Goose, Qwen Code, Crush) uses the same desk method and window as HM-SWE-Pro 400. Figures are conservative and labelled desk. They are not a live Scale board.

Open-index mix (kind, via, subscribe, trust) is counted from the listing catalog at render. Editorial install counts stay editorial. Unhoused slips (community bucket, generic MCP) are not assigned a fake studio.

ACES Skill Lift (arXiv:2608.20614) is paired with/without on a fixed task, harness, workspace, scorer. It is the closest published cousin of this desk’s skill ablation. Scan-only skill gates (Spearman ρ=0.14 vs live judges) are not a substitute.

IDWhoYearWhat we took
2608.26218Lewis2026Same Model, Different Harness. Tight-window Verified: F2PF 28→49, complete 43→72.
2608.20614Kevin et al. (NVIDIA ACES)2026Skill Lift 0.2134 on 947 paired cases. Scan-only vs judge ρ=0.14.
2608.14036Jiang et al. (UCSD)2026Skills as procedural anchors (65.7%). Retrieval precision 29.6%→3.3% as pool 5→100.
2608.22152Sun et al.2026Collaboration tax. 32 tasks, 11 models, 7 providers. EMNLP 2026.
2608.16801Aste2026File-mediated coordination −42% output tokens at 8 agents on message-heavy work.
2608.17188Shamay2026Six production token patterns. 60–70% cut. 2,420 confirmatory trials.
2601.14470Salim et al.2026Tokenomics. ChatDev+GPT-5, 30 tasks. Review 59.4% of tokens. Input 53.9%.
2606.17819Skill-eval @ KDD workshop2026500 skills, 1,000 tasks, 19 model configs. Instruction-following varies widely.
2606.11435Ding et al.2026Survey: skill evaluation and evolution. Six benchmark categories.
2602.12430Skills survey2026Architecture, acquisition, security. 26.1% of public skills flagged in empirical analyses.
2512.24565Liu et al. MCPAgentBench2026Real-world MCP tool-use with distractors in the candidate list.
2606.05304PACT2026Action-state communication. −38.7% tokens; SWE-agent input −50.4%.
2603.15183Parakhin2026Token Coherence / MESI for artifacts. Simulated 84–95% sync savings.
2605.06320LATTE2026Adaptive task graphs. Token cost 47.5% of next-best. Overwrites 4.3 vs 22.8–35.4.
2606.17454Gupta et al. SSA2026Simple Strands Agent. Minimal harness, SOTA-level Verified / Pro / TB2.
SEAL-ProScale SEAL2026SWE-bench Pro public, SWE-agent / mini-swe-agent. Live board ≠ paper-era ~23%.
IRIS-MHStanford IRIS Meta-Harness2026Opus 4.6 on TB2.0: 43% → 76.4% by harness only.
MCP-UMCP-Universe2026GPT-5-High 44.16% / Grok-4 33.33% on 11 live MCP servers.
NMI-MASNature Mach. Intell. / MIT2026MAS turn overhead 58–515%. SAS 67.7 successes/1k tok vs hybrid 13.6.
SkillfulSkillful.sh2026-08-29553,855 tools across 55 directories. Skills 58.2 / MCP 37.5 / agents 4.3.