The beginner series · 6 min

Tokens, Context Windows & Memory: Why Does My AI Forget Shit?

What a model can see now, and why “memory” is overloaded.

Tokens are the chunks a model reads and writes — roughly word-pieces, not whole words. The context window is the model's desk: everything it can see for one request, and it's finite. “Memory” is the slippery word — it's several different things wearing one name (conversation history replayed back in, facts an app stores for you, knowledge pulled in by RAG), none of which is the model quietly remembering you. Your AI “forgets” because the fact fell off the desk, not because it changed its mind.

Reading promise: No assumed AI knowledge. Jargon gets translated before it gets used.

First: What Is a Token?

Language models do not process text as a clean list of English words. Text is tokenized into commonly occurring character sequences. OpenAI’s documentation gives examples where a common word may be one token while a longer word can split into multiple pieces.[1]

Human sees:
"The dog ran through the backyard."

Model receives token IDs corresponding to chunks roughly like:
"The" | " dog" | " ran" | " through" | " the" | " backyard" | "."
Token ≠ word. A token is a chunk used by the model’s input/output representation.

Input Tokens, Output Tokens, Reasoning Tokens

YOU ── input tokens ──→ MODEL ── output tokens ──→ YOU
                         │
                some models also use
                   reasoning tokens

This matters because a model request has a finite context budget. OpenAI currently defines the context window as the maximum number of tokens usable in a single request, including input, output, and—where applicable—reasoning tokens.[2]

The Desk Analogy for Context

Imagine the model has a desk. During a request, the application places useful material on that desk:

┌────────────── CONTEXT WINDOW ──────────────┐
│ system/developer instructions              │
│ conversation history                       │
│ retrieved documents                        │
│ tool definitions                           │
│ tool results                               │
│ current question                           │
│ room for the answer                        │
└────────────────────────────────────────────┘

A larger context window gives the application a larger desk. It does not mean the model permanently remembers every conversation you have ever had.

Why Conversation History Can Feel Like Memory

Message 1: "My dog's name is Waffles."
Message 2: "What food should I buy?"
Message 3: "What's his name?"

If the application includes Message 1 again in the model input,
the model can simply read "Waffles" from the current context.

That is state or conversation history. It may feel like memory because the application keeps putting old material back on the desk.

Context Window ≠ Long-Term Memory

ThingUseful mental model
Context windowHow large the desk can be for one request
Current contextWhat is actually on the desk now
Conversation storeOld messages saved outside the immediate request
Long-term memoryFacts/state intentionally persisted and retrieved later
Model weightsPatterns learned during training
RAGA way to retrieve external knowledge into the current context

Several Things Get Called “Memory”

1. Model weights

Training changes parameters inside the model. This is not the same as a database of exact sentences, but it is one place learned patterns and capabilities are encoded.

2. Current context

Temporary working material the model can directly condition on during the current inference.

3. Conversation history

An application can store prior messages and supply some or all of them again.

4. Application memory

A product can deliberately persist structured facts or summaries and retrieve them when relevant.

5. External knowledge retrieval

RAG retrieves information from documents/databases and inserts it into context. It can behave like “memory” from the user’s perspective even though the information lives outside the model.

Everything Useful Eventually Has to Reach the Context

MEMORY DB ─┐
RAG INDEX ─┼─ retrieve/select ─→ CONTEXT ─→ LLM
MCP TOOL ──┤
FILES ─────┘

A brilliant memory system is useless if it never retrieves the right thing. Likewise, a huge database does not help the model until the application chooses something from it and supplies it to the model.

Context Gets Crowded—Fast

This becomes obvious with agents. A security investigation could accumulate an alert, 500 authentication events, 10,000 SIEM events, endpoint telemetry, email logs, DNS records, threat intelligence, and repeated tool retries. The agent’s context grows with each observation.

START              ███
AFTER SIEM          █████████
AFTER EDR           ██████████████
AFTER IDENTITY      ███████████████████
AFTER THREAT INTEL  █████████████████████████

OpenAI’s Codex-loop write-up calls out exactly this problem: repeated tool calls can grow prompt state until context management and compaction become necessary.[3]

Why Not Put Everything on the Desk?

Because can fit is not the same as is useful. Extra material can add cost, latency, contradictions, irrelevant instructions, stale observations, and distraction. Context engineering is the discipline of deciding what information should be present for the current decision.

A desk the size of Manhattan solves the space problem. It does not solve the “which of these 400,000 papers actually matters?” problem.

How Systems Manage Long Context

1. Drop old or low-value turns.

2. Summarize older material into a compact state.

3. Store details externally and retrieve them only when needed.

4. Use RAG rather than permanently stuffing entire knowledge bases into context.

5. Keep structured investigation/task state separate from raw logs.

OpenAI’s current compaction capability exists specifically to shrink long-running interaction state while preserving what later turns need.[4]

A Better Security-Agent Memory

Instead of keeping every raw log line, the agent can maintain compressed state and pointers back to evidence:

INVESTIGATION STATE

User: alice@example.com
Host: LAPTOP-4821

Key findings:
- unusual login from residential IP
- suspicious ZIP downloaded
- PowerShell executed 3 minutes later

Open questions:
- malicious payload?
- persistence?
- lateral movement?

Evidence pointers:
- Splunk search #482
- EDR event #98322
- email #11383

That is closer to how a skilled analyst works: keep the thesis and important facts in working memory, retain references to the raw evidence, and reopen detail when necessary.

When “The Model Forgot” Is the Wrong Diagnosis

  • The fact fell out of the supplied context.
  • The application intentionally truncated old turns.
  • A summary omitted the detail.
  • The memory system stored it but failed to retrieve it.
  • The model saw it but failed to use it correctly.
  • Newer conflicting information dominated the answer.
  • The application never persisted it at all.

The Cheat Sheet

TermPlain English
TokenA chunk the model processes
Context windowThe maximum working-space budget for a request
ContextWhat the model is actually given now
Conversation statePrior interaction data an application may carry forward
MemoryPersisted information that can be retrieved later
CompactionReduce accumulated context while preserving important state
Context engineeringChoose what should be on the desk
Tokens determine how much space information consumes. The context window limits the desk. Memory/RAG/tools decide what information can be brought back to the desk.

Sources

  • [3] OpenAI — Unrolling the Codex agent loop.