Essay · Agent Memory × Systems · Interactive

Your Context Window Is a Cache. Start Measuring It Like One.

Agent memory has been treated as a storage question — what to keep, where to put it. It is a caching question, and caching is a solved field with sixty years of measurement behind it. Borrowing three of its tools shows that most agents are leaving a threefold context improvement on the floor.

Series: Building agents that finish (part 11 of 11)
Also in this series: Memory · The Hippocampus · Memory Needs Forgetting · Context Compaction · Memory & Weights · The Memory Wall
The short version

A small fast store, backed by a big slow one, with a policy deciding what stays. That is a CPU cache, and it is also an agent’s context window sitting in front of its memory store. The architecture world has been measuring this exact structure since the 1960s, and three of its tools transfer directly.

Here is a structure you have seen before. A small, fast, expensive store holds what you are working on. Behind it sits a large, slow, cheap store holding everything else. When you need something that is not in the fast store you pay a large penalty to go and fetch it, and to make room you must throw something out. The quality of your whole system turns on the policy that decides what.

That is a CPU cache in front of DRAM. It is also, exactly, an agent’s context window in front of its memory store. We have been writing about the second as though it were a new problem — what to store, what to forget, what compaction destroys — while a field that has been optimizing the first since the 1960s sits one building over with the vocabulary already built.

01The mapping, which is tighter than an analogy

This is not a loose metaphor. Every structural element has a counterpart, and the counterparts behave the same way:

Computer architectureAgentShared behaviour
CacheThe context windowSmall, fast, expensive, fixed size
Main memory / diskMemory files, vector store, the repoLarge, slow, cheap, effectively unbounded
Cache lineA retrieved chunkThe unit you move, not the unit you need
Cache missA retrieval — or a step taken without the factCosts far more than a hit
Eviction policyCompaction policyDecides what survives when space runs out
PrefetchingSpeculative retrieval before the step needs itFree if idle bandwidth exists, harmful if wrong
Working setWhat this task actually needs in handIf it does not fit, everything thrashes
ThrashingThe agent re-reading the same file every few stepsAll traffic, no progress

The last row is worth dwelling on, because it names something every agent engineer has watched and few have had a word for. An agent that keeps re-reading the same file, re-running the same search, re-deriving the same fact is not confused. It is thrashing — its working set does not fit in its cache, so every access evicts something it is about to need. Architecture has had that diagnosis since 1968, along with the fix: either grow the cache or shrink the working set, and growing is usually the expensive one.

02The equation worth stealing

Cache people do not argue about whether a change helped; they compute average memory access time. For agents the units are tokens, seconds or dollars — pick one and stay consistent:

AMAT = hit_time + miss_rate × miss_penalty # hit_time the fact is already in context: near-free, but not free - # it occupies tokens and dilutes attention # miss_rate fraction of needed facts that were not in context # miss_penalty a retrieval round trip - or worse, a step executed on a # missing fact and the recovery that follows

The value is not the arithmetic, it is that it forces the comparison teams are currently making by instinct. Three ways to improve, and the equation tells you which is yours:

You could…Which term it movesWorth doing when
Use a bigger context windowmiss_rate ↓, hit_time ↑Misses are capacity misses. Note the second effect — a longer context costs more attention per step, so this is not free.
Improve retrieval qualitymiss_rate ↓You are fetching the wrong things, not too few.
Make retrieval cheaper or prefetchmiss_penalty ↓Misses are unavoidable but recovery is slow.
Fix the eviction policymiss_rate ↓Almost always — and almost nobody checks. The rest of this post is why.

03Three kinds of miss, and only one is about size

Hill and Smith’s classification splits cache misses into three kinds, and the split matters because the fixes are different:

MissIn a cacheIn an agentThe fix
CompulsoryFirst ever reference — unavoidableThe agent has never seen this file. It must read it once.Prefetching only. Cannot be removed.
CapacityThe working set does not fitThe task genuinely needs more in hand than the window holdsBigger context, or decompose the task
ConflictEvicted something it should have keptCompaction dropped the planA better policy. Costs nothing.

Compulsory misses set a floor: every distinct item must be fetched at least once, so no policy at any size beats it. Capacity misses are real and cost money to fix. Conflict misses are free money — the item was available, you had room for it at some point, and the policy threw it away anyway.

And here is the thing nobody measures: when an agent fails because it forgot the requirement, teams reach for a bigger context window, which is the fix for a capacity miss. If it was a conflict miss, that purchase buys very little, and the actual fix was free.

04The lab: four policies, one trace

Below is a simulated agent run — 300 accesses over 103 distinct items. Some are anchors (the plan, the constraints, the user’s goal) consulted throughout. Some are working files accessed in bursts. Some are one-shot tool outputs. Four eviction policies compete on the identical trace, including Belady’s optimal.

Start on LRU — least-recently-used, which is what a sliding context window implements and what most compaction schemes approximate. 38.7% hit rate at a context of 14. Now look at the trace strip underneath: the tall bars are accesses to the plan and constraints, and the white-capped ones are the times the agent went to use the plan and it was gone.

Now press LFU, which evicts by frequency instead of recency. Same trace, same context size, same everything: 53.7%. Fifteen points, for changing one line of policy.

Agent traces have inverted locality, and it breaks the default

Caches are built on a reliable assumption: what you touched recently you will touch again soon. For instruction and data streams that is overwhelmingly true, and it is why LRU has been the sensible default for decades.

Agent traces violate it in a specific way. The most important items have the longest gaps between uses. The plan is consulted at step 3, then step 40, then step 190 — important every time, never recent. Meanwhile the tool output from thirty seconds ago is maximally recent and will never be read again. Recency ranks these exactly backwards, which is why the plan is the first thing you forget: not a quirk of summarizers, a predictable consequence of ranking by recency on a workload whose value is uncorrelated with it.

05Belady’s optimal: the oracle you are allowed to compute

In 1966 László Bélády proved what the best possible eviction policy is: evict the item whose next use is furthest in the future. It is optimal, and it is unimplementable, because it requires knowing the future.

So it is usually taught as a curiosity. That is a mistake, because of one detail: after a run has finished, you have the future. You have the whole access trace. You can replay it, compute what Belady would have done, and compare. That turns an impossible algorithm into a measurement instrument:

conflict_misses ≈ misses(your_policy) − misses(OPT) # everything OPT also missed was compulsory or capacity - unavoidable. # the difference is purely your policy's fault, and it is free to fix.

Press OPT in the lab. At a context of 14 it reaches 61.3% against LRU’s 38.7% — so on this trace, more than half of LRU’s misses were nobody’s fault but the policy’s. No extra context, no better retriever, no larger model would have recovered them.

Now drag the context slider and watch where each curve flattens, because that is the result I find most useful:

PolicyHit rate at 14 slotsSlots needed to reach the floor
OPT (oracle)61.3%20
LFU (frequency)53.7%64
LRU (recency)38.7%69
FIFO (arrival)34.0%100

“The floor” is the compulsory-miss limit — 65.7% here — past which extra capacity buys nothing for any policy.

Your eviction policy is worth 3.45× your context window

Optimal eviction extracts everything a context of 20 can give. LRU needs 69 slots to reach the same place. On this trace, a better policy is worth about 3.45× the context — and context is the thing teams pay for, queue for, and build entire architectures around.

That ratio is a property of this trace, not a universal constant. The point is that it is measurable on yours, from a log you probably already have, in an afternoon.

› Go deeper: why LFU is not the answer either, and what to actually ship

LFU wins here because anchors are genuinely frequent, but pure frequency has a known pathology: cache pollution. An item that was hot early accumulates a count that keeps it resident long after it stops mattering, and new items cannot displace it. Real systems use aged or windowed counts (ARC, LIRS, and the TinyLFU family all exist because neither pure recency nor pure frequency is good enough).

The agent-shaped version of this is simpler than any of them, and it is the one I would ship: do not learn importance, declare it. Pin the task description, the constraints and the current plan so they are never eviction candidates at all, and run whatever policy you like over the remainder. In cache terms you are reserving a way for a known-hot line. It captures most of the gap in the table above without needing to predict anything, and it is a few lines of code.

Then measure the rest against OPT. If the remaining gap is small, your policy is fine and the misses are capacity misses, so go buy context. If the gap is large, keep tuning the policy, because context will not save you.

06Why this is the same wall as the hardware one

There is a reason cache theory transfers so cleanly here, and it is not that the analogy is pretty. It is that the hardware memory wall is physically underneath the agent one.

A long context is not free in the way "it fits, so it is fine" suggests. Every token of context becomes KV cache, which must be read from HBM on every subsequent token. For a 70B model that is 320 KiB per token of context, re-read for every token generated. Your agent’s "fast store" is fast only in comparison to a retrieval round trip — in absolute terms it is being dragged across a memory bus at 3.35 TB/s, and the bus is already the bottleneck.

# the same hierarchy, four levels deep SRAM → HBM FlashAttention tiles to avoid this trip HBM → host / disk KV offload and paging live here context → memory store retrieval and compaction live here context → weights what is learned vs what is looked up

Every level is the same bargain: a small fast tier, a large slow one, and a policy in between that decides what gets to be close. The reason your agent cannot simply remember everything is the same reason your GPU sits idle — moving bytes costs more than using them, at every scale in the stack, and the ratio has been getting worse for twenty years.

Which also explains why “just use a million-token context” disappoints in practice. It converts conflict misses into hits, yes — but it raises hit_time on every single step, because each of those tokens is re-read from HBM for every token generated. The cache equation already told you the tradeoff; the hardware decides the exchange rate.

07Recap

One sentence, with everything now behind it: an agent’s context is a cache, so measure it like one. Concretely, four things:

1. AMAT = hit_time + miss_rate × miss_penalty tells you which of the four fixes is actually yours 2. classify the misses compulsory = floor, capacity = buy context, conflict = free to fix 3. replay the trace against Belady's OPT the gap is exactly what your policy cost you 4. pin the anchors plan, constraints, goal: never eviction candidates

The field spent sixty years learning that the policy matters more than the size, that recency is a proxy and not a truth, and that the only honest way to evaluate a cache is against an oracle. Agent memory gets to skip the sixty years. It mostly has not yet.

Takeaways

  1. Context is a cache. Hit rate, AMAT, working set, thrashing, prefetch — the vocabulary already exists and it is more precise than the one agent engineering is using.
  2. Agent traces have inverted locality. The most valuable items have the longest reuse distances, which is the exact workload where recency-based eviction is worst.
  3. LFU beat LRU by 15 points on the same trace at the same size, because frequency tracks importance better than recency does here.
  4. Belady’s optimal is computable after the fact. Replay a logged run and the gap to OPT is precisely the portion of your misses that a free policy change would have prevented.
  5. A better policy was worth 3.45× the context window on this trace — 20 slots against 69 for the same hit rate.
  6. Pin the anchors. Most of the gap, a few lines of code, no prediction required.

Related reading

The hardware wall underneath this one · What compaction destroys, and why it compounds · Agent memory: what to store and where · Memory needs forgetting · Context or weights? · Four kinds of memory

On the numbers. The lab is a simulation over a stated synthetic trace — 300 accesses, 103 distinct items, a seeded generator mixing long-reuse anchors, bursty working files and one-shot outputs — not telemetry from a real agent. Every figure quoted above was computed independently before it was written and re-checked against the shipped widget, including the invariants that Belady’s optimal dominates every other policy at every capacity and that no policy beats the compulsory-miss floor. The 3.45× ratio is a property of this trace; what transfers is the method for measuring it on yours. Belady’s result dates to 1966; the three-kinds-of-miss classification to Hill and Smith.

Read next: Your GPU Is Not Computing. It Is Waiting. · The Plan Is the First Thing You Forget.

Cite this post
@misc{murugesan2026contextcache,
  author = {Murugesan, Sugeerth},
  title  = {Your Context Window Is a Cache. Start Measuring It Like One.},
  year   = {2026},
  month  = {oct},
  url    = {https://sugeerth.github.io/blog/context-as-cache/},
  note   = {Accessed: [date]}
}
SM
Sugeerth Murugesan Staff ML Engineer / Scientist · Intel / Intuit