Writing

Agent Infrastructure Is Becoming a Product Category

Every team used to build the loop, the store, the sandbox. That layer is being sold rather than written.

Notes

For the first wave of agent products, infrastructure was something every team built around the model themselves. Write an agent loop. Store the messages. Execute the tools. Create a sandbox if code needs to run. Add retries when tasks fail. Build context compression when sessions become too long.

That layer is starting to become a product of its own.

OpenAI's new Agents API packages the Codex harness, long-running session management, context compaction, tool routing, subagents and execution environments behind one managed API. Anthropic has been moving in the same direction with Managed Agents. AWS now has Bedrock AgentCore, with runtime, memory, identity, gateways, observability and persistent compute designed around agents rather than ordinary request-response applications.

I think this marks a category change. We are no longer only buying access to models. We are beginning to buy infrastructure designed for software whose control flow is partly generated at runtime.

The model API is no longer enough

A normal model API has a simple contract: send input, receive output.

An agent task can last hours or days. It may produce files, launch processes, call dozens of tools, delegate work to subagents and cross several context windows before it reaches a result.

Keeping that system reliable requires machinery outside the model.

OpenAI says the Agents API is built around the same harness and infrastructure used by Codex. It manages the agent loop while developers choose the task, model, tools and execution environment. Sessions can continue across context limits through compaction, while subagents can keep separate contexts and work in parallel.

Those are runtime services that used to live in application code.

The harness has become managed infrastructure

Earlier this year I wrote that the harness is part of capability. The same model can behave differently depending on context policy, tool design, filesystem access, memory and execution feedback.

The Agents API takes the next step: the harness itself is now something a provider can operate and update for you.

OpenAI describes the service as versioned access to an evolving Codex harness. As model behaviour changes, the provider can change context handling, tool exposure and multi-agent coordination without forcing every developer to rebuild those mechanisms.

Anthropic's Managed Agents follows a related idea by separating the session log, harness and sandbox. Harness assumptions can become stale as models improve, so stable interfaces around those layers matter.

This changes where product teams can spend their engineering effort.

Instead of implementing another generic agent loop, they can focus on the domain contract: which tools exist, what data the agent can access, what counts as completion and what actions require approval.

The execution environment is becoming a service boundary

The second part of this category is compute.

Long-running agents need a place to work. They need filesystems, packages, processes, temporary artifacts and sometimes persistent state across sessions.

OpenAI now offers hosted sandboxes but lets developers connect their own infrastructure or use third-party sandbox providers. The hosted environment can run code, work with files and produce artifacts while the harness remains managed separately.

AWS has been building a similar layer more explicitly with Bedrock AgentCore Runtime. Its serverless option isolates sessions in dedicated microVMs. Its newer runtime instances support persistent, multi-day sessions, GPU workloads and multiple collaborating agents on shared managed compute.

This looks much closer to cloud infrastructure than to a chatbot API.

The useful abstraction is no longer simply generate(). It is closer to: run this task in this environment, under these permissions, and preserve enough state for it to finish.

Agent infrastructure is splitting into recognizable subsystems

The stack is becoming easier to name.

There is a model for reasoning, a harness for the agent loop and context, an execution environment for commands and artifacts, durable state, identity and permissions, interoperability through protocols such as MCP or A2A, observability, and evaluation.

AWS AgentCore is almost a map of this decomposition. Its product surface includes Runtime, Memory, Identity, Gateway and Observability rather than one monolithic "agent" endpoint. OpenAI is packaging many of the same needs around the Codex harness, hosted sandboxes and durable sessions.

When several vendors independently start naming and selling the same infrastructure concerns, I usually take that as a sign that the abstraction is becoming real.

This resembles the evolution of ordinary cloud software

Early web applications bundled many concerns together. As the category matured, databases, queues, identity, storage, observability and compute became separate infrastructure products. Agents appear to be going through a compressed version of that process: custom loops first, frameworks next, managed infrastructure now.

I do not expect one vendor to own the whole stack. In fact, the current architecture points in the opposite direction.

OpenAI lets developers choose their execution environment. AgentCore is framework- and model-flexible and supports MCP and A2A. Anthropic's Managed Agents separates the harness from the sandbox. These systems are beginning to define boundaries where one layer can change without replacing everything around it.

That matters because models still move faster than infrastructure.

A production agent may change from one model generation to another several times while its permissions, data sources, completion checks and execution environment remain mostly stable.

What this changes for production builders

For founders, I would now make a deliberate build-versus-buy decision around the agent runtime.

If the product's value is the execution infrastructure itself, owning the harness may make sense. If the product's value is a domain workflow, rebuilding generic context compaction, sandbox lifecycle and subagent orchestration may be wasted engineering time.

The decision resembles cloud infrastructure choices more than prompt-engineering choices.

For creative production, I would keep the domain layer under my control. A Blender agent should still use project-specific validation rules; a game-production agent should still understand build states, asset ownership and publishing permissions. Generic infrastructure can handle sessions, compute and orchestration. Domain systems should own truth.

The downside is dependency on the runtime

Managed infrastructure creates another kind of lock-in.

If agent quality depends heavily on one provider's harness, moving the model may not reproduce the same behaviour somewhere else. Context policies, tool-search behaviour, subagent coordination and recovery logic can become hidden parts of application performance.

OpenAI partly addresses this by basing the Agents API on the open-source Codex harness. That helps, but managed services will still differ in execution semantics and security boundaries. I would keep domain tools and data interfaces portable even when the runtime is managed.

MCP helps at the tool boundary. Standard artifact formats help at the output boundary. External evaluation suites help verify that changing runtimes did not quietly change product behaviour.

The goal should not be zero dependency. It should be knowing where the dependency lives.

My prediction

I think agent infrastructure will settle into a recognizable software category in the same way model APIs did.

Teams will evaluate runtimes based on session durability, context management, sandbox options, tool interoperability, security, observability, concurrency, recovery and cost. Model support will remain one dimension rather than the whole product.

This could change AI application architecture in a useful way. The model becomes replaceable reasoning compute. The runtime manages execution. Domain services expose controlled capabilities. Evaluation decides whether the work succeeded.

Not every agent needs this machinery. A three-step workflow should remain a three-step workflow. Wrapping simple automation in a managed multi-agent runtime can create cost and complexity without improving the result.

But the moment a task can run for hours, cross context windows, manipulate files, spawn parallel work and survive infrastructure failures, we are no longer describing a prompt loop.

We are describing a workload.

And workloads eventually get infrastructure built specifically for them.

Sources

  1. OpenAI, Introducing the Agents API, 10 September 2026.
  2. OpenAI, The next evolution of the Agents SDK, 15 April 2026.
  3. Anthropic, Scaling Managed Agents: Decoupling the brain from the hands, 8 April 2026.
  4. AWS, Runtime instances: persistent compute for production AI agents on Amazon Bedrock AgentCore, 6 August 2026.
  5. AWS, Amazon Bedrock AgentCore Runtime documentation, 2026.

More

Other write-ups

15 September 2026 Approval Is Not Publication Thirty seven items in the queue, one approved, nothing published. The last arrow in the diagram is the only one that pays. 2 min read 14 September 2026 Neural Rendering Is Crossing From Reconstruction Into Synthesis Reconstruction filled in what sparse sampling missed. DLSS 5 generates appearance the renderer never computed. 6 min read 14 September 2026 The Renderer Is Becoming a Training Data Engine The renderer used to sit at the end of the pipeline. Physical AI gives the same scene a second job: teaching a model. 6 min read 13 September 2026 Local AI Is Becoming a Compute Fabric Local AI meant one model on one machine. Routing inference across the devices already on a network changes the unit. 6 min read 12 September 2026 Choosing a local model with a stopwatch, not a benchmark Three models, five hard tasks, code actually executed. The official build was 2.4 times faster and more accurate than a community repack of the same model. 3 min read 12 September 2026 We measured real time lighting against baked light, and baked won A 3D product simulation that had to look like an offline render. The real time version ran at 60 frames per second and looked like clay. Here is the measurement and the architecture that replaced it. 4 min read 12 September 2026 What actually broke in an agency run by agents Three failures from a delivery stack that runs on AI. None of them were the model's fault, and all three reported success while producing nothing. 4 min read 11 September 2026 A Task Without A Check Command Is Not Automated If a task has no command that can fail, the pipeline advances on the appearance of work. 2 min read 10 September 2026 A Gate The Model Writes Is A Gate The Model Loosens Three quality gates returned green while the work behind them was wrong, each for a different reason. 2 min read 8 September 2026 The Scoring Model Was Wrong And It Put The Worst Lead First A weighted sum let one axis substitute for the other, so a company with money and no problem ranked in the top twenty. 2 min read 5 September 2026 Building Software Got Easy. Getting Value Out Of It Did Not Aristo took weeks to build. Everything after the build is still in progress, and that gap is the whole story. 3 min read 4 September 2026 Capability Is Becoming an Operational Risk Surface Safety questions used to be about the text. Once a model can act, the capability itself becomes something to operate. 6 min read 24 July 2026 The Scene Graph Is Becoming an API for AI A scene graph exists for artists and software. Agents are becoming another consumer, and they need structure rather than pixels. 6 min read 23 July 2026 Animation Is Moving From Clips to Motion Priors Authored keyframes and blended clips are giving way to asking which constraints define acceptable motion. 6 min read 22 July 2026 Materials Are Becoming Learned Programs A material is texture maps, parameters and shader code. It is starting to become a small learned program that answers a rendering question. 6 min read 16 July 2026 Procedural Systems Are Expanding Beyond Geometry Geometry Nodes started with a narrow name. Blender 5.2 puts physics, sound and object data through the same graph. 6 min read 29 June 2026 The Unit of AI Work Is Becoming the Task, Not the Turn Chat taught us to think one turn at a time. Long-running agents make the task the thing that is scheduled, resumed and reviewed. 6 min read 24 June 2026 Game Engines Are Becoming Operating Systems for Worlds Engines have been judged on what they render and simulate. The Unreal 6 roadmap points at operating a world rather than drawing one. 6 min read 11 June 2026 The Model Is Becoming a Replaceable Backend Choosing a provider used to mean choosing an architecture. A stable interface makes replacement possible and evaluation makes it safe. 6 min read 17 April 2026 The Harness Is Part of the Capability The same model behaves differently depending on context policy, tool design and execution feedback. That surrounding software is not neutral. 6 min read 19 March 2026 Physics Engines Are Becoming Trainable Components A simulator predicts what happens next. A differentiable one can answer which parameter should change to stop the failure. 6 min read 13 March 2026 The Agent Needs an Environment, Not Just Tools A search function and a database query were enough for short loops. Longer work needs a place to stand. 6 min read 9 February 2026 Coding Agents Are Becoming General-Purpose Computer Workers Repositories were a friendly environment: text in, terminal actions, checkable results. That was a starting point, not a boundary. 6 min read 11 December 2025 Open Standards Outlive Model Generations A year after the MCP bet, the argument can be checked against what happened rather than what was hoped. 7 min read 20 November 2025 Colour Management Is a Pipeline Contract Blender 5.0 reads as better display options. Giving a file an explicit working colour space is an architectural change. 6 min read 11 August 2025 The Best Model May Be a Router, Not a Model GPT-5 moves model selection inside the system. The interesting unit stops being which model and becomes which compute policy. 6 min read 8 August 2025 World Models Are Not Game Engines Yet Genie 3 generates a navigable 720p world at 24 fps. Production work needs state you can inspect when something goes wrong. 6 min read 26 May 2025 Memory Is Becoming a System Capability A follow-up to the long-context argument. Storing, selecting and expiring facts is turning into a named part of the product. 6 min read 19 May 2025 Coding Agents Change the Unit of Software Work AI coding tools have been judged where code appears on screen. The boundary moves when the agent owns a task instead of a snippet. 6 min read 11 April 2025 Agents Need Protocols Between Each Other, Not Just Tools Tool calling solves the inside of the loop. It says nothing about one agent reaching another built by a different team on another platform. 7 min read 13 March 2025 Agents Need a Runtime, Not a Prompt Loop An agent is no longer well described as a prompt inside a while loop. The useful abstraction owns execution around the model. 6 min read 28 February 2025 Reasoning Is Not the Only Path to Better Models Longer thinking improves maths and code. A model that solves a logic puzzle and misreads ordinary intent is not the better production model. 6 min read 9 January 2025 Rendering Is Becoming a Reconstruction Stack Sparse samples, motion data and lower-resolution frames become a larger final result. Debugging becomes layered when reconstruction sits in the middle. 6 min read 16 December 2024 Agent Reliability Is an Evaluation Problem, Not a Prompting Problem When an agent misses a step, the usual fix is a stricter prompt. The failure is more often in how completion is detected. 6 min read 29 November 2024 MCP Might Matter More Than Another Model Release A model can reason well and still be useless inside a company if it cannot reach the files, repositories and tools where work lives. 6 min read 28 October 2024 Computer Use Is the Missing Layer Between Models and Software Most integrations assume useful software exposes the right API. Much of real software never did. 6 min read 19 September 2024 Inference-Time Compute Is a New Scaling Axis o1 improves when it is allowed to spend longer on a problem. A benchmark score without a compute budget is an incomplete number. 6 min read 15 August 2024 The Final Pixel Won't Come From the Renderer Geometry, camera and scene structure stay reliable ground truth. More of final appearance is moving into learned systems. 6 min read 29 July 2024 Open Models Are Becoming Research Infrastructure Llama 3.1 gets discussed as a benchmark result. The licence terms change which experiments are possible at all. 6 min read 24 June 2024 The Model Is Becoming a Runtime Function calling, code execution and structured output turn inference into a loop. The model stops being a text generator and starts being a control layer. 6 min read 16 May 2024 Multimodality Changes the Architecture, Not Just the Interface GPT-4o is easy to read as a faster interface. Training one model end to end across text, vision and audio is an architectural change. 7 min read 11 March 2024 Benchmark Scores Are Not Model Capability Claude 3 posts 86.8% on MMLU and 50.4% on GPQA Diamond. The chart is useful and it is not the same thing as capability. 6 min read 20 February 2024 Long Context Is Not Memory Gemini 1.5 makes a million tokens usable. A larger working set is not a system that decides what should survive the session. 6 min read