Writing

Inference-Time Compute Is a New Scaling Axis

o1 improves when it is allowed to spend longer on a problem. A benchmark score without a compute budget is an incomplete number.

Notes

For the last few years, the dominant scaling story in AI has been easy to summarise: larger models, more training data, more training compute.

OpenAI's o1 release adds another axis that deserves separate attention. OpenAI says o1 improves not only when more compute is spent during training, but when the model is allowed to spend more time working on a problem at inference. That means capability is no longer determined only by what happened before deployment. Part of it can be purchased per problem, at the moment the problem is being solved.

I think this changes how we should think about model size, latency and even product architecture.

Bigger models are not the only way to spend compute

Traditional scaling happens mostly before the user sends a prompt. A company trains a larger model, spends more compute on the training run, and ships the resulting checkpoint. At inference, the interaction is relatively fixed: prompt in, tokens out.

o1 makes that boundary less rigid.

OpenAI says the model was trained with reinforcement learning to use a chain of thought, refine strategies, recognise mistakes and try another approach when the current one fails. The company reports that performance rises with both train-time compute and test-time compute.

The benchmark results make the effect visible. On the 2024 AIME problems, OpenAI reports that its full o1 model averaged 74% with one sample, 83% when taking consensus over 64 samples, and 93% when re-ranking 1,000 samples with a learned scoring function.

That is the part I find more interesting than the leaderboard position.

The same model can produce different levels of performance depending on how much computation we are willing to spend after the prompt arrives.

This idea did not begin with o1

There have been signals for a while.

Self-consistency showed that instead of accepting one chain-of-thought sample, a system could generate several reasoning paths and select the answer they converge on. Google researchers reported large gains on several arithmetic and reasoning benchmarks from that change in decoding alone.

Tree of Thoughts pushed the idea further by letting a language model explore and evaluate multiple intermediate paths rather than committing to one left-to-right completion. On the Game of 24 task, the paper reports 74% success with its search method versus 4% for GPT-4 with ordinary chain-of-thought prompting.

Verifier-based work points in the same direction. Instead of trusting the first answer, generate candidates, score their reasoning, and spend extra computation deciding which path deserves to survive.

What o1 changes is that this is no longer only an external prompting or search trick around a general model. OpenAI is training a model family around the idea that more inference work should improve the answer.

That feels like a different scaling regime.

The August papers make the trade-off clearer

Two papers published last month make the point in a more explicit way.

Wu and colleagues study what they call inference scaling laws: how model size, generated tokens, voting and search methods trade against one another under a compute budget. Their experiments show cases where a smaller Llemma-7B model paired with a stronger inference strategy outperforms Llemma-34B using a simpler strategy on MATH.

Snell and colleagues reach a related result. Their paper studies ways to allocate test-time compute and finds that the best method depends on problem difficulty. With a compute-optimal strategy, they report more than a fourfold efficiency improvement over a basic best-of-N approach. In one FLOPs-matched setting, extra test-time computation lets a smaller model outperform a model 14 times larger on problems where the smaller model already has some chance of solving the task.

That last condition matters. More thinking is not magic. If the underlying model has no useful representation of the problem, sampling it repeatedly may only produce more wrong answers.

The useful idea is not "compute longer and everything improves." It is that model size and inference budget can trade against each other in some regimes.

My takeaway: capability becomes partly dynamic

This changes the mental model I use for inference.

Today we usually choose a model before a request arrives. Cheap model for simple work. Expensive model for hard work. The selected model then runs with roughly the same inference pattern every time.

I think that will become more dynamic.

A system should be able to inspect a task and decide how much computation it deserves. A trivial extraction request should not receive the same reasoning budget as a difficult algorithm problem. A localisation lookup should return quickly. A complicated build failure, shader bug or codebase refactor may justify a slower pass with search, verification or multiple candidate solutions.

Compute becomes something the application allocates, not only something the model provider spent during training.

That is a meaningful architectural change.

Production systems will need compute budgets

For developers, the obvious cost is latency.

A model that thinks longer before answering will be slower. If the extra work involves multiple candidate generations, search or verification, it consumes more inference resources as well. The product therefore needs a policy for when that trade is worth making.

I can imagine this being useful in production pipelines I already work with.

A Blender assistant changing one material parameter does not need an expensive reasoning loop. A system trying to diagnose why an automated render pipeline produces the wrong result across several scenes might. The second task can involve logs, scripts, render settings and dependencies. Spending more compute to compare hypotheses and verify a proposed fix makes sense if the failure is costly.

Game development has the same split. Renaming an asset is cheap. Investigating an intermittent build error across code, configuration and platform state is not.

This suggests that AI products may need something similar to quality settings in rendering: not one fixed mode, but a budget chosen for the task.

Fast answers when speed matters. More inference work when correctness is worth waiting for.

Evaluation changes with it

Once inference compute becomes variable, model evaluation gets harder.

A benchmark score without a compute budget is incomplete. One sample and 64 samples are not the same result, a verifier changes what the number means, and the count of candidate paths explored and the time the model received both belong next to the score.

OpenAI's own o1 results make this obvious. AIME performance changes sharply between single-sample evaluation, consensus and re-ranking at much larger sample counts.

For production teams, I think this means evaluation needs a cost axis next to accuracy.

The useful question is not simply, "Which model gets the highest score?"

It is closer to: how much quality can I buy for a given latency and inference budget on my actual task distribution?

That is a much more practical metric for software.

Where I think this goes

My expectation is that inference-time compute will become a first-class part of model serving over the next few years.

Models may learn to spend little computation on obvious requests and much more on tasks where uncertainty remains high. Systems around them may route work between direct generation, multiple candidates, verification and search depending on difficulty.

This could weaken the assumption that the strongest system always needs the largest base model. In some workloads, a smaller model with a well-designed inference process may be a better cost-performance choice than a much larger model used in a single pass.

There are limits. Search can waste compute. Verifiers can be wrong. Longer reasoning can follow a bad path for longer. Hard tasks will still benefit from stronger base models.

But o1 changes the scaling discussion for me.

The question is no longer only how much computation we can spend to create a model.

It is how much computation the system should spend on this problem, right now.

Sources

  1. OpenAI, Learning to reason with LLMs, 12 September 2024.
  2. Snell et al., Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters, 6 August 2024.
  3. Wu et al., Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models, 1 August 2024.
  4. Wang et al., Self-Consistency Improves Chain of Thought Reasoning in Language Models, ICLR 2023.
  5. Yao et al., Tree of Thoughts: Deliberate Problem Solving with Large Language Models, NeurIPS 2023.

More

Other write-ups

15 September 2026 Approval Is Not Publication Thirty seven items in the queue, one approved, nothing published. The last arrow in the diagram is the only one that pays. 2 min read 14 September 2026 Neural Rendering Is Crossing From Reconstruction Into Synthesis Reconstruction filled in what sparse sampling missed. DLSS 5 generates appearance the renderer never computed. 6 min read 14 September 2026 The Renderer Is Becoming a Training Data Engine The renderer used to sit at the end of the pipeline. Physical AI gives the same scene a second job: teaching a model. 6 min read 13 September 2026 Local AI Is Becoming a Compute Fabric Local AI meant one model on one machine. Routing inference across the devices already on a network changes the unit. 6 min read 12 September 2026 Agent Infrastructure Is Becoming a Product Category Every team used to build the loop, the store, the sandbox. That layer is being sold rather than written. 6 min read 12 September 2026 Choosing a local model with a stopwatch, not a benchmark Three models, five hard tasks, code actually executed. The official build was 2.4 times faster and more accurate than a community repack of the same model. 3 min read 12 September 2026 We measured real time lighting against baked light, and baked won A 3D product simulation that had to look like an offline render. The real time version ran at 60 frames per second and looked like clay. Here is the measurement and the architecture that replaced it. 4 min read 12 September 2026 What actually broke in an agency run by agents Three failures from a delivery stack that runs on AI. None of them were the model's fault, and all three reported success while producing nothing. 4 min read 11 September 2026 A Task Without A Check Command Is Not Automated If a task has no command that can fail, the pipeline advances on the appearance of work. 2 min read 10 September 2026 A Gate The Model Writes Is A Gate The Model Loosens Three quality gates returned green while the work behind them was wrong, each for a different reason. 2 min read 8 September 2026 The Scoring Model Was Wrong And It Put The Worst Lead First A weighted sum let one axis substitute for the other, so a company with money and no problem ranked in the top twenty. 2 min read 5 September 2026 Building Software Got Easy. Getting Value Out Of It Did Not Aristo took weeks to build. Everything after the build is still in progress, and that gap is the whole story. 3 min read 4 September 2026 Capability Is Becoming an Operational Risk Surface Safety questions used to be about the text. Once a model can act, the capability itself becomes something to operate. 6 min read 24 July 2026 The Scene Graph Is Becoming an API for AI A scene graph exists for artists and software. Agents are becoming another consumer, and they need structure rather than pixels. 6 min read 23 July 2026 Animation Is Moving From Clips to Motion Priors Authored keyframes and blended clips are giving way to asking which constraints define acceptable motion. 6 min read 22 July 2026 Materials Are Becoming Learned Programs A material is texture maps, parameters and shader code. It is starting to become a small learned program that answers a rendering question. 6 min read 16 July 2026 Procedural Systems Are Expanding Beyond Geometry Geometry Nodes started with a narrow name. Blender 5.2 puts physics, sound and object data through the same graph. 6 min read 29 June 2026 The Unit of AI Work Is Becoming the Task, Not the Turn Chat taught us to think one turn at a time. Long-running agents make the task the thing that is scheduled, resumed and reviewed. 6 min read 24 June 2026 Game Engines Are Becoming Operating Systems for Worlds Engines have been judged on what they render and simulate. The Unreal 6 roadmap points at operating a world rather than drawing one. 6 min read 11 June 2026 The Model Is Becoming a Replaceable Backend Choosing a provider used to mean choosing an architecture. A stable interface makes replacement possible and evaluation makes it safe. 6 min read 17 April 2026 The Harness Is Part of the Capability The same model behaves differently depending on context policy, tool design and execution feedback. That surrounding software is not neutral. 6 min read 19 March 2026 Physics Engines Are Becoming Trainable Components A simulator predicts what happens next. A differentiable one can answer which parameter should change to stop the failure. 6 min read 13 March 2026 The Agent Needs an Environment, Not Just Tools A search function and a database query were enough for short loops. Longer work needs a place to stand. 6 min read 9 February 2026 Coding Agents Are Becoming General-Purpose Computer Workers Repositories were a friendly environment: text in, terminal actions, checkable results. That was a starting point, not a boundary. 6 min read 11 December 2025 Open Standards Outlive Model Generations A year after the MCP bet, the argument can be checked against what happened rather than what was hoped. 7 min read 20 November 2025 Colour Management Is a Pipeline Contract Blender 5.0 reads as better display options. Giving a file an explicit working colour space is an architectural change. 6 min read 11 August 2025 The Best Model May Be a Router, Not a Model GPT-5 moves model selection inside the system. The interesting unit stops being which model and becomes which compute policy. 6 min read 8 August 2025 World Models Are Not Game Engines Yet Genie 3 generates a navigable 720p world at 24 fps. Production work needs state you can inspect when something goes wrong. 6 min read 26 May 2025 Memory Is Becoming a System Capability A follow-up to the long-context argument. Storing, selecting and expiring facts is turning into a named part of the product. 6 min read 19 May 2025 Coding Agents Change the Unit of Software Work AI coding tools have been judged where code appears on screen. The boundary moves when the agent owns a task instead of a snippet. 6 min read 11 April 2025 Agents Need Protocols Between Each Other, Not Just Tools Tool calling solves the inside of the loop. It says nothing about one agent reaching another built by a different team on another platform. 7 min read 13 March 2025 Agents Need a Runtime, Not a Prompt Loop An agent is no longer well described as a prompt inside a while loop. The useful abstraction owns execution around the model. 6 min read 28 February 2025 Reasoning Is Not the Only Path to Better Models Longer thinking improves maths and code. A model that solves a logic puzzle and misreads ordinary intent is not the better production model. 6 min read 9 January 2025 Rendering Is Becoming a Reconstruction Stack Sparse samples, motion data and lower-resolution frames become a larger final result. Debugging becomes layered when reconstruction sits in the middle. 6 min read 16 December 2024 Agent Reliability Is an Evaluation Problem, Not a Prompting Problem When an agent misses a step, the usual fix is a stricter prompt. The failure is more often in how completion is detected. 6 min read 29 November 2024 MCP Might Matter More Than Another Model Release A model can reason well and still be useless inside a company if it cannot reach the files, repositories and tools where work lives. 6 min read 28 October 2024 Computer Use Is the Missing Layer Between Models and Software Most integrations assume useful software exposes the right API. Much of real software never did. 6 min read 15 August 2024 The Final Pixel Won't Come From the Renderer Geometry, camera and scene structure stay reliable ground truth. More of final appearance is moving into learned systems. 6 min read 29 July 2024 Open Models Are Becoming Research Infrastructure Llama 3.1 gets discussed as a benchmark result. The licence terms change which experiments are possible at all. 6 min read 24 June 2024 The Model Is Becoming a Runtime Function calling, code execution and structured output turn inference into a loop. The model stops being a text generator and starts being a control layer. 6 min read 16 May 2024 Multimodality Changes the Architecture, Not Just the Interface GPT-4o is easy to read as a faster interface. Training one model end to end across text, vision and audio is an architectural change. 7 min read 11 March 2024 Benchmark Scores Are Not Model Capability Claude 3 posts 86.8% on MMLU and 50.4% on GPQA Diamond. The chart is useful and it is not the same thing as capability. 6 min read 20 February 2024 Long Context Is Not Memory Gemini 1.5 makes a million tokens usable. A larger working set is not a system that decides what should survive the session. 6 min read