Writing

The Harness Is Part of the Capability

The same model behaves differently depending on context policy, tool design and execution feedback. That surrounding software is not neutral.

Notes

A year ago, it was still common to talk about agent performance as if it were mostly a property of the model. Give a stronger model the same prompt and tools, and the system should become more capable.

That picture is getting harder to defend.

OpenAI's April 15 update to the Agents SDK introduces what it calls a model-native harness: configurable memory, sandbox-aware orchestration, filesystem tools, MCP, skills, AGENTS.md instructions, shell execution and file editing around the model. OpenAI's own description is revealing: the harness is meant to get more out of a frontier model by matching the way those models perform best.

The model still matters. But I think the surrounding harness is becoming part of what we mean by capability.

Same model, different system

An agent benchmark does not test a naked model.

Something decides the system prompt. Something exposes tools. Something manages context when the task becomes long. Something determines how command output is returned, how many turns the model receives and when the agent is allowed to stop.

That software is the scaffold or harness.

Epoch AI published a useful measurement of this in December. On SWE-bench Verified, simply changing the scaffold produced differences of up to 11 percentage points for GPT-5 and up to 15 points for Kimi K2 Thinking. Epoch concluded that scaffold choice had the largest effect on overall performance in its setup.

This makes a benchmark score harder to interpret as a pure model property.

If the same checkpoint receives better tools, cleaner feedback and a context policy suited to its behaviour, the resulting agent can solve more tasks without any change to the weights. A poor harness can hide capability that the model already has.

The harness is where model intelligence meets software

The new Agents SDK makes this boundary unusually explicit.

OpenAI separates the harness from the sandbox where model-generated commands execute. The harness manages the agent loop, memory and orchestration. The sandbox owns files, commands and dependencies. A Manifest describes what the workspace should contain so the same agent logic can run against different execution providers.

The harness can expose MCP tools, skills, project instructions, shell commands and file-editing operations. It can shape where inputs appear, where outputs should be written and how work survives across long tasks.

These choices change what the model sees and what actions are easy for it to take.

The model is the reasoning engine. The harness decides how that reasoning touches state.

Context policy is part of performance

Context management is one of the clearest examples.

Long-running agents eventually produce more history than should stay active in the context window. The system has to compress, reset, retrieve or externalise some of it.

There is no universally correct policy.

Anthropic's work on long-running agents shows how model-specific this can become. In earlier experiments, Claude Sonnet 4.5 tended to rush toward completion as its context window filled. Anthropic added context resets and handoff artifacts to keep the work moving across clean sessions. When the same design was used with Opus 4.5, Anthropic found that the behaviour had disappeared and the reset logic had become unnecessary overhead.

That is a useful warning.

Harnesses encode assumptions about model failure modes. As models change, those assumptions can become wrong.

The best harness is not one giant collection of agent tricks. It is the smallest execution structure that matches the behaviour of the model and task.

Memory and tools shape apparent intelligence

The same applies to memory.

A model can have access to persistent storage and still use it badly. The system needs rules around what belongs in active context, what should be written to durable storage, and what should be discarded.

Anthropic's earlier long-running-agent work used progress files, feature lists, Git history and handoff artifacts to bridge context windows. Those mechanisms do more than increase storage. They shape attention: a progress file can stop repeated work, while a feature list can stop an agent from declaring victory too early.

Tool design has a similar effect.

A model given unrestricted shell access can theoretically perform thousands of operations. That does not mean unrestricted shell access is the best interface for every task.

A narrow tool can encode domain knowledge in its contract. validate_asset() is easier to use reliably than asking the model to rediscover every Blender validation rule through Python each time. run_tests() can return a structured summary instead of twenty thousand lines of terminal output.

A good harness reduces unnecessary decisions.

This matters because agent failure can come from interface friction rather than lack of reasoning. The model may understand what needs to happen but struggle to locate the file, parse noisy output, preserve state between steps or determine whether an action succeeded.

Security and durability live outside the model

As agents gain broader execution access, the harness becomes a security and recovery boundary.

OpenAI's updated architecture separates the control layer from the compute environment. Credentials can stay outside the sandbox where model-generated code executes. Agent state can be externalised so a failed container can be discarded and restored from a snapshot rather than treated as irreplaceable.

Anthropic has reached a related architecture in Managed Agents. It separates the session log, harness and sandbox so each can fail or change independently. Anthropic describes the harness as the loop that calls Claude and routes tool calls, while the sandbox performs computation and file operations.

The agent can request an operation. Infrastructure should decide whether it is allowed, where it runs and what state survives afterward.

What this changes in production

For a production system, I would now evaluate the model and harness separately.

The first measurement is whether the model can reason about the task at all.

The second is whether the harness makes the required state visible, exposes the right operations, preserves progress, verifies completion and recovers from failure.

In a Blender pipeline, one harness might give the model raw Python and the entire project directory. Another might expose asset metadata, narrow scene operations, render validation and an isolated workspace with explicit output paths.

Both systems can use the same model. I would expect very different reliability.

The same applies to game development. A model working through a noisy terminal with no project instructions is not equivalent to the same model with repository conventions, deterministic build commands, test summaries, task state and clean checkpoints.

Once agents operate for hours rather than seconds, those differences accumulate.

Benchmarking needs to name the harness

This changes how I read agent benchmark results.

A useful report should identify the model, but it should name the scaffold too. I want to know the tools, context strategy, execution environment, turn budget, reasoning settings and completion checks.

Epoch AI makes the distinction explicit: a standardised scaffold is useful when the goal is comparing models, while a model-native product harness may be better for measuring the frontier capability of the complete system.

Those are different questions.

"Which model is stronger under the same interface?" is a model evaluation.

"How much work can this agent system complete?" is a system evaluation.

Mixing the two produces misleading comparisons.

My prediction

I think harness engineering will become its own part of AI infrastructure.

Model providers will keep improving checkpoints, but production teams will compete on how well they expose state, manage context, design tools, isolate execution and verify results around those models.

The harness will need to evolve with the model. Some techniques that improve today's agent may become unnecessary when the next model handles the same problem internally. Other problems will move outward into infrastructure because software can solve them more cheaply and predictably than inference.

That means capability will become harder to summarise with a model name.

A deployed agent's real capability will increasingly be the capability of the whole stack:

model + context policy + tools + memory + environment + verification.

The weights set the ceiling.

The harness determines how much of that ceiling the system can actually reach.

Sources

  1. OpenAI, The next evolution of the Agents SDK, 15 April 2026.
  2. Anthropic, Scaling Managed Agents: Decoupling the brain from the hands, 8 April 2026.
  3. Anthropic, Harness design for long-running application development, 24 March 2026.
  4. Anthropic, Effective harnesses for long-running agents, 26 November 2025.
  5. Epoch AI, Why benchmarking is hard, 23 December 2025.

More

Other write-ups

15 September 2026 Approval Is Not Publication Thirty seven items in the queue, one approved, nothing published. The last arrow in the diagram is the only one that pays. 2 min read 14 September 2026 Neural Rendering Is Crossing From Reconstruction Into Synthesis Reconstruction filled in what sparse sampling missed. DLSS 5 generates appearance the renderer never computed. 6 min read 14 September 2026 The Renderer Is Becoming a Training Data Engine The renderer used to sit at the end of the pipeline. Physical AI gives the same scene a second job: teaching a model. 6 min read 13 September 2026 Local AI Is Becoming a Compute Fabric Local AI meant one model on one machine. Routing inference across the devices already on a network changes the unit. 6 min read 12 September 2026 Agent Infrastructure Is Becoming a Product Category Every team used to build the loop, the store, the sandbox. That layer is being sold rather than written. 6 min read 12 September 2026 Choosing a local model with a stopwatch, not a benchmark Three models, five hard tasks, code actually executed. The official build was 2.4 times faster and more accurate than a community repack of the same model. 3 min read 12 September 2026 We measured real time lighting against baked light, and baked won A 3D product simulation that had to look like an offline render. The real time version ran at 60 frames per second and looked like clay. Here is the measurement and the architecture that replaced it. 4 min read 12 September 2026 What actually broke in an agency run by agents Three failures from a delivery stack that runs on AI. None of them were the model's fault, and all three reported success while producing nothing. 4 min read 11 September 2026 A Task Without A Check Command Is Not Automated If a task has no command that can fail, the pipeline advances on the appearance of work. 2 min read 10 September 2026 A Gate The Model Writes Is A Gate The Model Loosens Three quality gates returned green while the work behind them was wrong, each for a different reason. 2 min read 8 September 2026 The Scoring Model Was Wrong And It Put The Worst Lead First A weighted sum let one axis substitute for the other, so a company with money and no problem ranked in the top twenty. 2 min read 5 September 2026 Building Software Got Easy. Getting Value Out Of It Did Not Aristo took weeks to build. Everything after the build is still in progress, and that gap is the whole story. 3 min read 4 September 2026 Capability Is Becoming an Operational Risk Surface Safety questions used to be about the text. Once a model can act, the capability itself becomes something to operate. 6 min read 24 July 2026 The Scene Graph Is Becoming an API for AI A scene graph exists for artists and software. Agents are becoming another consumer, and they need structure rather than pixels. 6 min read 23 July 2026 Animation Is Moving From Clips to Motion Priors Authored keyframes and blended clips are giving way to asking which constraints define acceptable motion. 6 min read 22 July 2026 Materials Are Becoming Learned Programs A material is texture maps, parameters and shader code. It is starting to become a small learned program that answers a rendering question. 6 min read 16 July 2026 Procedural Systems Are Expanding Beyond Geometry Geometry Nodes started with a narrow name. Blender 5.2 puts physics, sound and object data through the same graph. 6 min read 29 June 2026 The Unit of AI Work Is Becoming the Task, Not the Turn Chat taught us to think one turn at a time. Long-running agents make the task the thing that is scheduled, resumed and reviewed. 6 min read 24 June 2026 Game Engines Are Becoming Operating Systems for Worlds Engines have been judged on what they render and simulate. The Unreal 6 roadmap points at operating a world rather than drawing one. 6 min read 11 June 2026 The Model Is Becoming a Replaceable Backend Choosing a provider used to mean choosing an architecture. A stable interface makes replacement possible and evaluation makes it safe. 6 min read 19 March 2026 Physics Engines Are Becoming Trainable Components A simulator predicts what happens next. A differentiable one can answer which parameter should change to stop the failure. 6 min read 13 March 2026 The Agent Needs an Environment, Not Just Tools A search function and a database query were enough for short loops. Longer work needs a place to stand. 6 min read 9 February 2026 Coding Agents Are Becoming General-Purpose Computer Workers Repositories were a friendly environment: text in, terminal actions, checkable results. That was a starting point, not a boundary. 6 min read 11 December 2025 Open Standards Outlive Model Generations A year after the MCP bet, the argument can be checked against what happened rather than what was hoped. 7 min read 20 November 2025 Colour Management Is a Pipeline Contract Blender 5.0 reads as better display options. Giving a file an explicit working colour space is an architectural change. 6 min read 11 August 2025 The Best Model May Be a Router, Not a Model GPT-5 moves model selection inside the system. The interesting unit stops being which model and becomes which compute policy. 6 min read 8 August 2025 World Models Are Not Game Engines Yet Genie 3 generates a navigable 720p world at 24 fps. Production work needs state you can inspect when something goes wrong. 6 min read 26 May 2025 Memory Is Becoming a System Capability A follow-up to the long-context argument. Storing, selecting and expiring facts is turning into a named part of the product. 6 min read 19 May 2025 Coding Agents Change the Unit of Software Work AI coding tools have been judged where code appears on screen. The boundary moves when the agent owns a task instead of a snippet. 6 min read 11 April 2025 Agents Need Protocols Between Each Other, Not Just Tools Tool calling solves the inside of the loop. It says nothing about one agent reaching another built by a different team on another platform. 7 min read 13 March 2025 Agents Need a Runtime, Not a Prompt Loop An agent is no longer well described as a prompt inside a while loop. The useful abstraction owns execution around the model. 6 min read 28 February 2025 Reasoning Is Not the Only Path to Better Models Longer thinking improves maths and code. A model that solves a logic puzzle and misreads ordinary intent is not the better production model. 6 min read 9 January 2025 Rendering Is Becoming a Reconstruction Stack Sparse samples, motion data and lower-resolution frames become a larger final result. Debugging becomes layered when reconstruction sits in the middle. 6 min read 16 December 2024 Agent Reliability Is an Evaluation Problem, Not a Prompting Problem When an agent misses a step, the usual fix is a stricter prompt. The failure is more often in how completion is detected. 6 min read 29 November 2024 MCP Might Matter More Than Another Model Release A model can reason well and still be useless inside a company if it cannot reach the files, repositories and tools where work lives. 6 min read 28 October 2024 Computer Use Is the Missing Layer Between Models and Software Most integrations assume useful software exposes the right API. Much of real software never did. 6 min read 19 September 2024 Inference-Time Compute Is a New Scaling Axis o1 improves when it is allowed to spend longer on a problem. A benchmark score without a compute budget is an incomplete number. 6 min read 15 August 2024 The Final Pixel Won't Come From the Renderer Geometry, camera and scene structure stay reliable ground truth. More of final appearance is moving into learned systems. 6 min read 29 July 2024 Open Models Are Becoming Research Infrastructure Llama 3.1 gets discussed as a benchmark result. The licence terms change which experiments are possible at all. 6 min read 24 June 2024 The Model Is Becoming a Runtime Function calling, code execution and structured output turn inference into a loop. The model stops being a text generator and starts being a control layer. 6 min read 16 May 2024 Multimodality Changes the Architecture, Not Just the Interface GPT-4o is easy to read as a faster interface. Training one model end to end across text, vision and audio is an architectural change. 7 min read 11 March 2024 Benchmark Scores Are Not Model Capability Claude 3 posts 86.8% on MMLU and 50.4% on GPQA Diamond. The chart is useful and it is not the same thing as capability. 6 min read 20 February 2024 Long Context Is Not Memory Gemini 1.5 makes a million tokens usable. A larger working set is not a system that decides what should survive the session. 6 min read