Writing

The Best Model May Be a Router, Not a Model

GPT-5 moves model selection inside the system. The interesting unit stops being which model and becomes which compute policy.

Notes

For most of the LLM era, model selection has been a user or developer decision. Pick the fast model for ordinary work. Pick the reasoning model when the task is harder. Pick a smaller model when cost matters.

GPT-5 changes that interaction in ChatGPT. OpenAI describes it as a unified system made from a fast model, a deeper reasoning model, and a real-time router that decides which one should handle the request. The router looks at conversation type, complexity, tool needs and explicit user intent.

That architecture interests me more than the name GPT-5.

The strongest AI product may not be one model that handles every request the same way. It may be a system that decides how much intelligence, latency and compute a task deserves before choosing how to solve it.

Model selection is becoming a runtime decision

The model picker made sense when the difference between models was easy to understand.

GPT-4o was fast and general. The o-series spent more inference compute on hard reasoning. A user could choose based on the task.

But that choice requires the user to diagnose the problem before the system solves it.

OpenAI is moving that decision inward. In ChatGPT, GPT-5 routes ordinary requests to a faster model and harder ones to GPT-5 Thinking. OpenAI says the router can use signals including task complexity, tool requirements and instructions such as asking the system to "think hard." The company says the router is trained from signals such as user model switches, preference data and measured correctness.

The implication is simple: selecting the computational strategy is becoming part of inference itself.

That is a different problem from making one checkpoint larger.

Compute should follow task difficulty

Last year, o1 made inference-time compute visible as a scaling axis. A harder problem could receive more computation after the prompt arrived instead of depending only on a larger pretrained model.

GPT-5 turns that idea into product architecture.

Not every request benefits equally from extra reasoning. OpenAI's developer guidance says GPT-5 supports minimal, low, medium and high reasoning effort. It specifically notes that increasing reasoning beyond low adds little to some simple long-context retrieval tasks, while harder visual reasoning tasks gain more.

That is exactly the trade I would expect production systems to make.

If I ask an assistant to rename a group of exported files according to an existing rule, I do not need maximum reasoning. If I ask it to diagnose a rendering failure that crosses Blender scripts, asset state, render settings and external tools, spending more compute may be justified.

The useful unit is no longer "which model do I use?"

It becomes "what compute policy does this task need?"

Routing can include more than reasoning depth

The GPT-5 implementation makes routing visible mainly as a choice between fast responses and deeper reasoning, but I think the broader architecture extends further.

A production system can route across several dimensions:

  • model size;
  • reasoning budget;
  • tool access;
  • latency target;
  • cost ceiling;
  • local versus hosted execution;
  • specialist versus general model.

The decision does not have to happen once.

A task could begin cheaply, discover that it is harder than expected, then escalate. A fast model could classify or prepare the task before a stronger model handles the difficult section. A reasoning model could be reserved for the small part of the workflow where additional inference compute changes the result.

This resembles ordinary compute scheduling more than traditional chatbot design.

We already do this elsewhere in software. Render farms assign different hardware or quality settings depending on the job. Build systems skip work that does not need to be repeated. Game engines allocate expensive effects selectively rather than running the maximum-quality path for every pixel.

AI inference is starting to need the same discipline.

One unified product can hide several computational systems

There is an important distinction between GPT-5 in ChatGPT and GPT-5 in the API.

OpenAI says the ChatGPT product is the routed system: fast model, reasoning model and router. In the API, gpt-5 refers to the reasoning model used for maximum performance, with separate gpt-5-mini and gpt-5-nano variants and configurable reasoning effort. The non-reasoning model used by ChatGPT is exposed separately as gpt-5-chat-latest.

That distinction matters because "GPT-5" is no longer one simple computational object across every surface.

From the user's perspective, a single product name can represent several execution paths.

I suspect this becomes normal.

People should not need a mental map of model families to ask a question. The application should know enough about cost, latency and task difficulty to choose a sensible path, while still exposing controls when the user or developer needs them.

Routing creates its own failure mode

Automatic routing removes complexity from the user, but it does not remove the decision.

It moves the decision into infrastructure.

If the router underestimates a task, the system can answer quickly and badly. If it overestimates everything, latency and cost rise without useful quality gains. Two superficially similar prompts may take different paths, which can make behaviour harder to reproduce.

For developers, that means the router itself needs evaluation.

I would want to know how often simple tasks are escalated unnecessarily, how often difficult tasks remain on the fast path, and whether routing changes tool-use reliability. Evaluation should measure the complete system, not only each underlying model in isolation.

This connects directly to the agent reliability problem. Once a system contains models, tools, memory, runtimes and routers, "model quality" explains only part of the observed behaviour.

The orchestration policy becomes part of capability.

Production systems should optimise a frontier, not a score

The usual benchmark question is: which model scores highest?

A routed system has a different optimisation target.

For each task distribution, I care about the relationship between quality, latency and cost. There may be no single best operating point.

A product-facing assistant may prefer a fast answer for most requests and escalate only when confidence is low. A coding agent modifying a repository may accept slower responses because a bad change costs more than additional inference. A real-time creative tool may have a strict latency budget and use deeper reasoning only for background operations.

GPT-5's API design reflects part of this idea. Alongside three model sizes, developers can vary reasoning effort and output verbosity. OpenAI explicitly recommends experimenting with the trade between performance, cost and latency for the actual use case rather than assuming the maximum setting is always best.

That is how I would design model infrastructure now.

Not "use the smartest model everywhere."

Use the cheapest execution path that still meets the quality requirement, then escalate when the task gives you evidence that it deserves more.

My prediction

I think model routing will become a larger part of AI system design than model selection.

Users will increasingly interact with one logical assistant while the infrastructure underneath chooses between several computational paths. Developers may still pin models when reproducibility or compliance requires it, but general-purpose products will benefit from dynamic allocation.

The interesting systems will not only route between model names. They will route between levels of reasoning, specialists, tools and execution environments.

Over time, this could make the question "which model are you using?" less informative.

A better question may be: what policy decides how this task gets solved?

GPT-5 is still a model family, and the router does not replace the intelligence of the models underneath it. A weak model does not become strong because a router selected it correctly.

But once several strong models offer different cost, speed and reasoning profiles, choosing among them becomes its own source of system performance.

The best model may still matter.

The best product may be the one that knows when not to use it.

Sources

  1. OpenAI, Introducing GPT-5, 7 August 2025.
  2. OpenAI, Introducing GPT-5 for developers, 7 August 2025.
  3. OpenAI, GPT-5 System Card, 7 August 2025.

More

Other write-ups

15 September 2026 Approval Is Not Publication Thirty seven items in the queue, one approved, nothing published. The last arrow in the diagram is the only one that pays. 2 min read 14 September 2026 Neural Rendering Is Crossing From Reconstruction Into Synthesis Reconstruction filled in what sparse sampling missed. DLSS 5 generates appearance the renderer never computed. 6 min read 14 September 2026 The Renderer Is Becoming a Training Data Engine The renderer used to sit at the end of the pipeline. Physical AI gives the same scene a second job: teaching a model. 6 min read 13 September 2026 Local AI Is Becoming a Compute Fabric Local AI meant one model on one machine. Routing inference across the devices already on a network changes the unit. 6 min read 12 September 2026 Agent Infrastructure Is Becoming a Product Category Every team used to build the loop, the store, the sandbox. That layer is being sold rather than written. 6 min read 12 September 2026 Choosing a local model with a stopwatch, not a benchmark Three models, five hard tasks, code actually executed. The official build was 2.4 times faster and more accurate than a community repack of the same model. 3 min read 12 September 2026 We measured real time lighting against baked light, and baked won A 3D product simulation that had to look like an offline render. The real time version ran at 60 frames per second and looked like clay. Here is the measurement and the architecture that replaced it. 4 min read 12 September 2026 What actually broke in an agency run by agents Three failures from a delivery stack that runs on AI. None of them were the model's fault, and all three reported success while producing nothing. 4 min read 11 September 2026 A Task Without A Check Command Is Not Automated If a task has no command that can fail, the pipeline advances on the appearance of work. 2 min read 10 September 2026 A Gate The Model Writes Is A Gate The Model Loosens Three quality gates returned green while the work behind them was wrong, each for a different reason. 2 min read 8 September 2026 The Scoring Model Was Wrong And It Put The Worst Lead First A weighted sum let one axis substitute for the other, so a company with money and no problem ranked in the top twenty. 2 min read 5 September 2026 Building Software Got Easy. Getting Value Out Of It Did Not Aristo took weeks to build. Everything after the build is still in progress, and that gap is the whole story. 3 min read 4 September 2026 Capability Is Becoming an Operational Risk Surface Safety questions used to be about the text. Once a model can act, the capability itself becomes something to operate. 6 min read 24 July 2026 The Scene Graph Is Becoming an API for AI A scene graph exists for artists and software. Agents are becoming another consumer, and they need structure rather than pixels. 6 min read 23 July 2026 Animation Is Moving From Clips to Motion Priors Authored keyframes and blended clips are giving way to asking which constraints define acceptable motion. 6 min read 22 July 2026 Materials Are Becoming Learned Programs A material is texture maps, parameters and shader code. It is starting to become a small learned program that answers a rendering question. 6 min read 16 July 2026 Procedural Systems Are Expanding Beyond Geometry Geometry Nodes started with a narrow name. Blender 5.2 puts physics, sound and object data through the same graph. 6 min read 29 June 2026 The Unit of AI Work Is Becoming the Task, Not the Turn Chat taught us to think one turn at a time. Long-running agents make the task the thing that is scheduled, resumed and reviewed. 6 min read 24 June 2026 Game Engines Are Becoming Operating Systems for Worlds Engines have been judged on what they render and simulate. The Unreal 6 roadmap points at operating a world rather than drawing one. 6 min read 11 June 2026 The Model Is Becoming a Replaceable Backend Choosing a provider used to mean choosing an architecture. A stable interface makes replacement possible and evaluation makes it safe. 6 min read 17 April 2026 The Harness Is Part of the Capability The same model behaves differently depending on context policy, tool design and execution feedback. That surrounding software is not neutral. 6 min read 19 March 2026 Physics Engines Are Becoming Trainable Components A simulator predicts what happens next. A differentiable one can answer which parameter should change to stop the failure. 6 min read 13 March 2026 The Agent Needs an Environment, Not Just Tools A search function and a database query were enough for short loops. Longer work needs a place to stand. 6 min read 9 February 2026 Coding Agents Are Becoming General-Purpose Computer Workers Repositories were a friendly environment: text in, terminal actions, checkable results. That was a starting point, not a boundary. 6 min read 11 December 2025 Open Standards Outlive Model Generations A year after the MCP bet, the argument can be checked against what happened rather than what was hoped. 7 min read 20 November 2025 Colour Management Is a Pipeline Contract Blender 5.0 reads as better display options. Giving a file an explicit working colour space is an architectural change. 6 min read 8 August 2025 World Models Are Not Game Engines Yet Genie 3 generates a navigable 720p world at 24 fps. Production work needs state you can inspect when something goes wrong. 6 min read 26 May 2025 Memory Is Becoming a System Capability A follow-up to the long-context argument. Storing, selecting and expiring facts is turning into a named part of the product. 6 min read 19 May 2025 Coding Agents Change the Unit of Software Work AI coding tools have been judged where code appears on screen. The boundary moves when the agent owns a task instead of a snippet. 6 min read 11 April 2025 Agents Need Protocols Between Each Other, Not Just Tools Tool calling solves the inside of the loop. It says nothing about one agent reaching another built by a different team on another platform. 7 min read 13 March 2025 Agents Need a Runtime, Not a Prompt Loop An agent is no longer well described as a prompt inside a while loop. The useful abstraction owns execution around the model. 6 min read 28 February 2025 Reasoning Is Not the Only Path to Better Models Longer thinking improves maths and code. A model that solves a logic puzzle and misreads ordinary intent is not the better production model. 6 min read 9 January 2025 Rendering Is Becoming a Reconstruction Stack Sparse samples, motion data and lower-resolution frames become a larger final result. Debugging becomes layered when reconstruction sits in the middle. 6 min read 16 December 2024 Agent Reliability Is an Evaluation Problem, Not a Prompting Problem When an agent misses a step, the usual fix is a stricter prompt. The failure is more often in how completion is detected. 6 min read 29 November 2024 MCP Might Matter More Than Another Model Release A model can reason well and still be useless inside a company if it cannot reach the files, repositories and tools where work lives. 6 min read 28 October 2024 Computer Use Is the Missing Layer Between Models and Software Most integrations assume useful software exposes the right API. Much of real software never did. 6 min read 19 September 2024 Inference-Time Compute Is a New Scaling Axis o1 improves when it is allowed to spend longer on a problem. A benchmark score without a compute budget is an incomplete number. 6 min read 15 August 2024 The Final Pixel Won't Come From the Renderer Geometry, camera and scene structure stay reliable ground truth. More of final appearance is moving into learned systems. 6 min read 29 July 2024 Open Models Are Becoming Research Infrastructure Llama 3.1 gets discussed as a benchmark result. The licence terms change which experiments are possible at all. 6 min read 24 June 2024 The Model Is Becoming a Runtime Function calling, code execution and structured output turn inference into a loop. The model stops being a text generator and starts being a control layer. 6 min read 16 May 2024 Multimodality Changes the Architecture, Not Just the Interface GPT-4o is easy to read as a faster interface. Training one model end to end across text, vision and audio is an architectural change. 7 min read 11 March 2024 Benchmark Scores Are Not Model Capability Claude 3 posts 86.8% on MMLU and 50.4% on GPQA Diamond. The chart is useful and it is not the same thing as capability. 6 min read 20 February 2024 Long Context Is Not Memory Gemini 1.5 makes a million tokens usable. A larger working set is not a system that decides what should survive the session. 6 min read