Foundational Mechanics · 7 min

The Context Window Is Not Memory

A model re-reads the whole transcript every turn from a finite, costly, unevenly-used window — and mistaking that for memory breeds production bugs.

Ask a model your name. Tell it. Start a fresh chat and ask again. It has no idea. That lands as a surprise because the thing clearly feels like it remembers: inside one conversation it tracks what you said ten messages back, corrects itself, picks up threads. So where did the memory go?

There was never any memory. What you felt was the context window: a block of tokens the model re-reads from scratch on every single turn. Getting this one distinction straight prevents more early production bugs than almost anything else you'll learn.

The mental model: a stateless function over text

At inference time a transformer is a pure function. Hand it a sequence of tokens, get back a probability distribution over the next token. That is the entire contract. Nothing carries over between calls. When a chat "remembers" your earlier message, it is because the application layer stitched the whole transcript back together and fed it in again as one long string:

Turn 3, what the model actually receives:
┌──────────────────────────────────────────────┐
│ [system prompt] You are a helpful assistant.  │
│ [user]      My name is Dan.                   │  ← re-sent
│ [assistant] Nice to meet you, Dan.            │  ← re-sent
│ [user]      What's my name?                   │  ← the "new" part
└──────────────────────────────────────────────┘
        all of this = the context window

The memory is just concatenation. Drop a message from that reconstructed transcript and the model's "memory" of it doesn't degrade, it vanishes. Nothing was persisted. That is why the field keeps a hard line between context (whatever sits in the window this turn) and memory (anything durable: a database, a vector store, a file the app re-reads and pastes back in).

The window has a ceiling, measured in tokens. Tokens aren't words. Each one runs about 4 characters of English, so roughly 750 words per 1,000 tokens. A model sold as "128K context" sees around 96,000 words at once, prompt and reply combined. Go over and something has to be cut, usually the oldest turns. Hence the second surprise people hit: the model didn't forget, the app quietly truncated.

Why the window is expensive

Attention lets every token look at every other token. For a sequence of length N that is an N×N comparison, so a full forward pass is quadratic in the sequence length. Double the context and you quadruple the attention work.

Generation claws some of that back with the KV cache. Each token's key and value vectors don't change once computed, so they get cached and reused, and producing each new token costs O(N) instead of reprocessing the entire prefix. The bill comes due as memory. The cache grows with every token in the window, and it lives in scarce GPU VRAM.

Concrete numbers for Llama-2-7B (32 layers, hidden size 4096, fp16):

KV per token = 2 (K and V) × 32 layers × 4096 dims × 2 bytes
             ≈ 512 KB / token

  4,096 tokens →  ~2 GB VRAM
 32,768 tokens → ~16 GB VRAM   (more than some models' entire weights)

Long context is not free headroom to fill "just in case." Every token you park in the prompt is latency and money on every turn it survives, and it fights the weights for the same VRAM. Grouped-query attention, paged attention, quantized caches: these shrink the constant, not the scaling.

Builder: Prompt cost scales with the whole window and gets re-billed each turn, because the whole thing is reprocessed each turn. That 20K-token system prompt you're proud of taxes message #1 and message #50 equally. Trim hard, and measure tokens, not characters.

Failure mode: lost in the middle

Here is the part that catches strong engineers. Even when the relevant fact fits comfortably inside the window, the model may not use it well, and where you put it changes the outcome a lot.

Liu et al. (2024, TACL) tested this head-on. They ran multi-document QA and key-value retrieval, then slid the single relevant document to different spots in a long context. Accuracy traced a U: strongest when the needed fact sat at the very start or the very end, sagging in the middle. The gap is not cosmetic. It runs into double digits of accuracy points, and in their harder settings a model with the answer buried mid-context did no better than the same model handed no documents at all. This held even for models marketed as long-context.

accuracy
  ▲
  │ ●                                       ●
  │   ●                                   ●
  │      ●                             ●
  │          ●                    ●
  │               ●___________●         ← "lost in the middle"
  └───────────────────────────────────────────►
   start        position of relevant info       end

So a big window is a budget, not a guarantee. Stuffing in 50 documents and trusting the model to find the right one is a bet you lose in the middle. Put the load-bearing content at the edges, most reliably near the end, right up against the question.

It helps to keep three separate failures from collapsing into one vague "long context is bad," because each has a different fix. Truncation is a resource ceiling — the tokens fell out of the window (above), so the answer is to spend the budget more carefully. Lost-in-the-middle is attention dilution — the tokens are present but neglected, so the answer is placement. And a third, distinct from both, is positional degradation: past the sequence length a model was trained on, the positional encoding itself starts to misbehave — the extrapolation problem from Module 5 (positional encodings). Same phrase, three mechanisms, three fixes; don't reach for placement when the real problem is that you blew past the trained context length.

Defender: Treat this as attack surface. If an adversary can get text into something the model will read (a support ticket, a scraped page, a résumé), they can drop a prompt-injection payload where attention is strongest, the tail, and shove your real instructions into the neglected middle. "It was in the context" is not "the model obeyed it." Window inclusion is never enforcement.
Researcher: The U-shape is a window into positional encoding and attention-head behavior, not a law of physics. Follow-up work ties it to how RoPE-style encodings (Module 5) and particular heads spend their weight, and newer long-context models narrow the dip without closing it. RULER and HELMET exist precisely because needle-in-a-haystack pass rates oversell real long-context reasoning.

A 15-minute lab

You can watch every claim above on a laptop model:

  1. Prove statelessness. In one API call, send "My name is Dan." and "What's my name?" as two messages. It answers. Now make two independent calls with the first message left out. It can't. The only thing that changed is what you concatenated.
  2. Watch truncation. Paste a 200-line document, ask about line 3, then keep chatting until the transcript passes the window limit. Ask again. The answer decays not because the model "forgot" but because your client evicted the oldest tokens. Log the token count each turn and you'll catch the exact moment it happens.
  3. Reproduce lost-in-the-middle. Take 20 short paragraphs, hide one fact ("the launch code is 4417"), and ask for it with the fact at position 1, then 10, then 20. Same tokens, same model, three different hit rates. Position 10 is the one that hurts.

Where this goes next

Once the window reads as finite, costly, and unevenly used, the shape of every serious LLM system follows. You don't dump everything in and pray. You retrieve the few chunks that matter — found by the same "embedding distance means similarity" geometry from Module 2 (embeddings) — place them on purpose, and bolt on the durable memory the model lacks: a database it can query, tools it can call, a retrieval step that decides what earns a slot in that expensive, attention-starved window.

That retrieval discipline is Module 4 (RAG), and it exists for exactly the reason this lesson opened with. RAG is just Module 2's similarity search wired to the front of the prompt: it picks the handful of chunks whose vectors sit closest to your question and pastes only those into the window. The context window is not memory, so we build the memory around it.

Sources

  • Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the Middle: How Language Models Use Long Contexts. Transactions of the ACL, 12, 157–173. https://aclanthology.org/2024.tacl-1.9/ (arXiv:2307.03172)
  • Vaswani, A., et al. (2017). Attention Is All You Need — self-attention and its O(N²) cost. arXiv:1706.03762
  • Kwon, W., et al. (2023). Efficient Memory Management for LLM Serving with PagedAttention (vLLM) — KV-cache memory as the serving bottleneck. arXiv:2309.06180
  • Hsieh, C.-P., et al. (2024). RULER: What's the Real Context Size of Your Long-Context Language Models? arXiv:2404.06654