HarnessmarketEnter the desk
/HMX Claude Code 29.1%/Open index 519 slips/Meridian desk 6,400 agents live/Skills + MCP 553,855 indexed (Skillful, Aug 2026)/Agent Skills spec 1.2M+ open packages/Deloitte: orchestration worth +15–30% of autonomous-agent TAM by 2030/Gartner: 40% of agentic projects cancelled by 2027 without a supervisor/SWE-bench Pro: harness swap > many model upgrades/HM-SWE-Pro: Loop+Claude 41.2 pass@1 · chat 18.2 · same model/Lewis 2608.26218: F2PF 28→49 under a tighter harness, same model/ACES: Skill Lift 0.21 · 947 paired cases · scan vs live ρ=0.14/Tokenomics: code review 59.4% of ChatDev tokens/HMX Claude Code 29.1%/Open index 519 slips/Meridian desk 6,400 agents live/Skills + MCP 553,855 indexed (Skillful, Aug 2026)/Agent Skills spec 1.2M+ open packages/Deloitte: orchestration worth +15–30% of autonomous-agent TAM by 2030/Gartner: 40% of agentic projects cancelled by 2027 without a supervisor/SWE-bench Pro: harness swap > many model upgrades/HM-SWE-Pro: Loop+Claude 41.2 pass@1 · chat 18.2 · same model/Lewis 2608.26218: F2PF 28→49 under a tighter harness, same model/ACES: Skill Lift 0.21 · 947 paired cases · scan vs live ρ=0.14/Tokenomics: code review 59.4% of ChatDev tokens

Thesis · 2026-08-24 · 14 min

The harness is the product

Models commoditise. The software shell that turns them into agents does not. 2026 is the year the industry started saying this out loud.

Harnessmarket Intelligence

In February 2026 the phrase 'harness engineering' left the labs. OpenAI's Ryan Lopopolo used it. Mitchell Hashimoto argued for it. Practitioners had already been living it: on SWE-bench Pro, swapping the agent harness moved pass@1 more than many model upgrades. That sentence is the founding fact of this market.

A harness is not a prompt. It is the loop that offers tools, the permission system that decides what those tools may touch, the memory that survives a session, the recovery that turns a failed call into a retry instead of a bill, and the stop condition that a model will not apply to itself. Think of the model as a CPU, the context window as RAM, and the harness as the operating system. Nobody prices an OS like a CPU cycle. The 2024–2025 market did, which is why so many 'agent companies' were model wrappers with a landing page.

The commercial pattern that remains is unromantic. Open source with a paid host. A vertical harness for a regulated industry. A marketplace of skills that the harness knows how to load. A control plane that will still answer the phone at ten thousand agents. Harnessmarket is the last two, and the registry that makes the first two tradable.

Three implications follow. First, procurement will move from 'which model' to 'which harness, which mix, which license'. Second, evaluation has to hold the model constant — Lattice's SWE Eval Pack exists because almost nobody does. Third, attribution is no longer a nice dashboard. If two harnesses touched the same ticket, legal, finance, and engineering all need a number. Token share is not that number.