Writing

Long Context Is Not Memory

Gemini 1.5 makes a million tokens usable. A larger working set is not a system that decides what should survive the session.

Notes

Google's Gemini 1.5 announcement made one number hard to ignore: one million tokens of context. The model ships with a standard 128K context window, while a limited group of developers and enterprise customers can test up to 1M tokens in private preview. Google says the model can find an inserted piece of information 99% of the time in a one-million-token "needle in a haystack" test.

That is a large step for working with codebases, long documents, video and mixed media. But I think the more useful distinction is what it does not solve.

A million-token context window is not the same thing as memory.

It gives a model a much larger working set. Memory is a system for deciding what should survive, how it should be stored, when it should be updated, and how it should return later. Those are different engineering problems.

Capacity is not persistence

It is easy to describe a long context window as memory because the model can "remember" something from much earlier in the prompt. Google itself uses that language carefully, describing long context as helping a model recall information during a session.

The phrase "during a session" matters.

Context is input state. If I place a design document, codebase and chat history into a prompt, the model can reason over that material while it remains inside the active context. If the next session does not contain the same material, there is no guarantee that the model has retained it. Nothing about a larger context window automatically creates a durable record.

Persistent memory needs at least three extra operations: write, select and retrieve. In a real system, it needs deletion, conflict handling, versioning and some idea of relevance as well.

This is closer to a database or storage hierarchy than to a larger prompt.

Retrieval, attention and context pollution are different problems

The easiest way to see the distinction is to separate retrieval from attention.

Retrieval-augmented generation puts an external search step in front of the model. A retriever selects a small set of documents or passages from a larger corpus, then sends those results into the model. The model never needs to see the whole corpus at once. The retrieval layer makes a relevance decision before generation begins.

A long-context model changes that balance. Instead of filtering aggressively before inference, we can place far more source material directly into the context and let the model resolve relevance inside its own computation.

Gemini 1.5 makes that approach much more practical. Google reports that 1.5 Pro can process around 700,000 words, an hour of video, 11 hours of audio, or large codebases inside a single prompt. That means fewer artificial boundaries between chunks of a project.

But "the model can see it" and "the model will use it correctly" are still separate claims.

Research published before Gemini 1.5 already showed why. The 2023 paper Lost in the Middle found that long-context language models could become less reliable depending on where relevant information appeared. Performance was often strongest when the needed information was near the beginning or end of the input, and weaker when it appeared in the middle.

Gemini 1.5's reported needle-retrieval results suggest a major improvement on this type of access. Yet needle retrieval is a narrow test: one piece of information is placed among distractors and the model is asked to recover it. Real project context is messier. It contains duplicated requirements, obsolete decisions, conflicting versions, half-finished notes and information that was once relevant but no longer is.

That creates what I would call context pollution. A bigger window reduces the need to throw information away, but it increases the amount of information the model must discriminate between.

The engineering problem moves from "How do I fit enough context?" to "What deserves to occupy the model's attention?"

Memory needs policy, not just capacity

The most interesting work may happen one layer above the model.

A useful memory system needs a policy for what becomes durable. A conversation contains many facts, but only a small subset should survive for future work. Some facts expire. Some should be attached to a project, not a person. Some should be replaced when a newer version exists. Some should never be stored at all.

Long context has no built-in answer to these questions. It only increases how much state can be presented at once.

MemGPT, published in 2023, is interesting here because it treats the problem more like an operating system. Instead of assuming the model should hold everything in one context, it proposes multiple memory tiers and moves information between them. The model gets a constrained active context while older or less relevant information can live in external storage and return when needed.

That architecture feels closer to where production systems need to go.

For me, the useful mental model is not "context versus retrieval" as competing approaches. It is a memory hierarchy:

  • durable storage for source-of-truth information;
  • retrieval for selecting relevant material;
  • active context as working memory;
  • the model's attention for resolving relationships inside that working set.

A one-million-token window makes the working-memory tier much larger. It does not remove the other tiers.

What this changes in production

This distinction matters more when AI is connected to real creative and technical workflows.

Take a Blender or game-development project. There may be asset naming rules, shader conventions, render settings, engine constraints, previous technical decisions, bug history, reference images and automation scripts. With a huge context window, it becomes tempting to send the entire project history into every request.

I would resist that design.

The better system is likely to keep durable project state outside the model, retrieve the parts related to the current task, then use long context when the task genuinely benefits from broad visibility. For example, reviewing a large codebase or comparing several versions of a production document may justify a wide context. Remembering that a project uses a specific coordinate convention or export rule should not require resending months of history.

This matters for cost and latency, but the larger issue is control. External memory can be inspected, edited, versioned and tested. A giant prompt is harder to reason about as a system boundary.

For founders building AI products, I would treat memory as infrastructure rather than a model feature. Store durable facts explicitly. Keep provenance. Track which memory was retrieved for a response. Separate user preferences from project state. Give old information a way to expire or be replaced.

Long context makes all of this easier because retrieval can be less brittle. We can fetch larger, more coherent blocks instead of slicing everything into tiny chunks. But it does not make memory architecture optional.

Where I think this is going

Context windows will probably keep growing. The value is obvious: fewer chunking artifacts, better codebase-level reasoning, longer multimodal inputs and richer in-context learning.

I do not expect that growth to eliminate retrieval or persistent storage. I expect the opposite: larger contexts will make the boundaries between storage, retrieval and inference more explicit.

As models become capable of consuming much larger working sets, application developers will have to decide what enters that set and why. The useful systems will probably combine long context with persistent, structured memory rather than choosing one or the other.

There is a useful analogy with 3D production. More VRAM lets me keep more geometry and textures available to the renderer, but it does not replace asset management. Capacity reduces one constraint. It does not decide which asset is current, which version is approved, or how the pipeline should reconstruct a scene tomorrow.

I think AI systems will end up with a similar separation.

The context window will become a larger workspace. Memory will remain the layer that decides what survives beyond it.

Sources

  1. Google, Our next-generation model: Gemini 1.5, 15 February 2024.
  2. Google, What is a long context window?, 16 February 2024.
  3. Liu et al., Lost in the Middle: How Language Models Use Long Contexts, 2023.
  4. Packer et al., MemGPT: Towards LLMs as Operating Systems, 2023.
  5. Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, 2020.

More

Other write-ups

15 September 2026 Approval Is Not Publication Thirty seven items in the queue, one approved, nothing published. The last arrow in the diagram is the only one that pays. 2 min read 14 September 2026 Neural Rendering Is Crossing From Reconstruction Into Synthesis Reconstruction filled in what sparse sampling missed. DLSS 5 generates appearance the renderer never computed. 6 min read 14 September 2026 The Renderer Is Becoming a Training Data Engine The renderer used to sit at the end of the pipeline. Physical AI gives the same scene a second job: teaching a model. 6 min read 13 September 2026 Local AI Is Becoming a Compute Fabric Local AI meant one model on one machine. Routing inference across the devices already on a network changes the unit. 6 min read 12 September 2026 Agent Infrastructure Is Becoming a Product Category Every team used to build the loop, the store, the sandbox. That layer is being sold rather than written. 6 min read 12 September 2026 Choosing a local model with a stopwatch, not a benchmark Three models, five hard tasks, code actually executed. The official build was 2.4 times faster and more accurate than a community repack of the same model. 3 min read 12 September 2026 We measured real time lighting against baked light, and baked won A 3D product simulation that had to look like an offline render. The real time version ran at 60 frames per second and looked like clay. Here is the measurement and the architecture that replaced it. 4 min read 12 September 2026 What actually broke in an agency run by agents Three failures from a delivery stack that runs on AI. None of them were the model's fault, and all three reported success while producing nothing. 4 min read 11 September 2026 A Task Without A Check Command Is Not Automated If a task has no command that can fail, the pipeline advances on the appearance of work. 2 min read 10 September 2026 A Gate The Model Writes Is A Gate The Model Loosens Three quality gates returned green while the work behind them was wrong, each for a different reason. 2 min read 8 September 2026 The Scoring Model Was Wrong And It Put The Worst Lead First A weighted sum let one axis substitute for the other, so a company with money and no problem ranked in the top twenty. 2 min read 5 September 2026 Building Software Got Easy. Getting Value Out Of It Did Not Aristo took weeks to build. Everything after the build is still in progress, and that gap is the whole story. 3 min read 4 September 2026 Capability Is Becoming an Operational Risk Surface Safety questions used to be about the text. Once a model can act, the capability itself becomes something to operate. 6 min read 24 July 2026 The Scene Graph Is Becoming an API for AI A scene graph exists for artists and software. Agents are becoming another consumer, and they need structure rather than pixels. 6 min read 23 July 2026 Animation Is Moving From Clips to Motion Priors Authored keyframes and blended clips are giving way to asking which constraints define acceptable motion. 6 min read 22 July 2026 Materials Are Becoming Learned Programs A material is texture maps, parameters and shader code. It is starting to become a small learned program that answers a rendering question. 6 min read 16 July 2026 Procedural Systems Are Expanding Beyond Geometry Geometry Nodes started with a narrow name. Blender 5.2 puts physics, sound and object data through the same graph. 6 min read 29 June 2026 The Unit of AI Work Is Becoming the Task, Not the Turn Chat taught us to think one turn at a time. Long-running agents make the task the thing that is scheduled, resumed and reviewed. 6 min read 24 June 2026 Game Engines Are Becoming Operating Systems for Worlds Engines have been judged on what they render and simulate. The Unreal 6 roadmap points at operating a world rather than drawing one. 6 min read 11 June 2026 The Model Is Becoming a Replaceable Backend Choosing a provider used to mean choosing an architecture. A stable interface makes replacement possible and evaluation makes it safe. 6 min read 17 April 2026 The Harness Is Part of the Capability The same model behaves differently depending on context policy, tool design and execution feedback. That surrounding software is not neutral. 6 min read 19 March 2026 Physics Engines Are Becoming Trainable Components A simulator predicts what happens next. A differentiable one can answer which parameter should change to stop the failure. 6 min read 13 March 2026 The Agent Needs an Environment, Not Just Tools A search function and a database query were enough for short loops. Longer work needs a place to stand. 6 min read 9 February 2026 Coding Agents Are Becoming General-Purpose Computer Workers Repositories were a friendly environment: text in, terminal actions, checkable results. That was a starting point, not a boundary. 6 min read 11 December 2025 Open Standards Outlive Model Generations A year after the MCP bet, the argument can be checked against what happened rather than what was hoped. 7 min read 20 November 2025 Colour Management Is a Pipeline Contract Blender 5.0 reads as better display options. Giving a file an explicit working colour space is an architectural change. 6 min read 11 August 2025 The Best Model May Be a Router, Not a Model GPT-5 moves model selection inside the system. The interesting unit stops being which model and becomes which compute policy. 6 min read 8 August 2025 World Models Are Not Game Engines Yet Genie 3 generates a navigable 720p world at 24 fps. Production work needs state you can inspect when something goes wrong. 6 min read 26 May 2025 Memory Is Becoming a System Capability A follow-up to the long-context argument. Storing, selecting and expiring facts is turning into a named part of the product. 6 min read 19 May 2025 Coding Agents Change the Unit of Software Work AI coding tools have been judged where code appears on screen. The boundary moves when the agent owns a task instead of a snippet. 6 min read 11 April 2025 Agents Need Protocols Between Each Other, Not Just Tools Tool calling solves the inside of the loop. It says nothing about one agent reaching another built by a different team on another platform. 7 min read 13 March 2025 Agents Need a Runtime, Not a Prompt Loop An agent is no longer well described as a prompt inside a while loop. The useful abstraction owns execution around the model. 6 min read 28 February 2025 Reasoning Is Not the Only Path to Better Models Longer thinking improves maths and code. A model that solves a logic puzzle and misreads ordinary intent is not the better production model. 6 min read 9 January 2025 Rendering Is Becoming a Reconstruction Stack Sparse samples, motion data and lower-resolution frames become a larger final result. Debugging becomes layered when reconstruction sits in the middle. 6 min read 16 December 2024 Agent Reliability Is an Evaluation Problem, Not a Prompting Problem When an agent misses a step, the usual fix is a stricter prompt. The failure is more often in how completion is detected. 6 min read 29 November 2024 MCP Might Matter More Than Another Model Release A model can reason well and still be useless inside a company if it cannot reach the files, repositories and tools where work lives. 6 min read 28 October 2024 Computer Use Is the Missing Layer Between Models and Software Most integrations assume useful software exposes the right API. Much of real software never did. 6 min read 19 September 2024 Inference-Time Compute Is a New Scaling Axis o1 improves when it is allowed to spend longer on a problem. A benchmark score without a compute budget is an incomplete number. 6 min read 15 August 2024 The Final Pixel Won't Come From the Renderer Geometry, camera and scene structure stay reliable ground truth. More of final appearance is moving into learned systems. 6 min read 29 July 2024 Open Models Are Becoming Research Infrastructure Llama 3.1 gets discussed as a benchmark result. The licence terms change which experiments are possible at all. 6 min read 24 June 2024 The Model Is Becoming a Runtime Function calling, code execution and structured output turn inference into a loop. The model stops being a text generator and starts being a control layer. 6 min read 16 May 2024 Multimodality Changes the Architecture, Not Just the Interface GPT-4o is easy to read as a faster interface. Training one model end to end across text, vision and audio is an architectural change. 7 min read 11 March 2024 Benchmark Scores Are Not Model Capability Claude 3 posts 86.8% on MMLU and 50.4% on GPQA Diamond. The chart is useful and it is not the same thing as capability. 6 min read