Emerging Frontiers · 7 min

Test-Time Learning and External Memory

"Infinite context" is a headline, not a mental model. This lesson replaces it with the honest version: ultra-long-context, external memory, and models that adapt their own state as they read — plus the six limits that survive even an endless stream.

"Infinite context" is a marketing slide. What actually ships is a number: 128K tokens, 200K, a million, ten million. Every one of those is finite, every one degrades before you hit the ceiling, and every one charges you — in latency and dollars — for tokens you paid to stuff in and then largely ignored. Build a system on the belief that you can keep pasting more text until the model "knows everything," and you will ship something that gets slower, more expensive, and quietly less accurate as it fills up. The honest frontier is more interesting than the fantasy. Here is what is actually being built to make a model remember things it wasn't trained on.

The three families, and why they exist

There are three distinct answers to "how does a model use information it sees at run time," and confusing them is the source of most bad architecture decisions.

Ultra-long-context keeps everything in the attention window. You extend the context length and let self-attention look at all of it. Simple, exact, and expensive: standard attention cost grows with the square of sequence length, so a 10x longer prompt is roughly 100x the attention compute unless you change the mechanism. Sparse or sliding-window attention (Longformer, BigBird) and better positional schemes (RoPE, plus length-extrapolation tricks like YaRN) push the ceiling higher, but the token still sits in a window someone has to store and scan.

External memory keeps information outside the model and fetches the relevant slice on demand. This is retrieval-augmented generation: embed your documents, store the vectors, and at query time pull the top-k chunks into the prompt. The model's weights never change; the memory is a database you own. This is what almost everyone should reach for first, and we'll come back to why.

Test-time adaptation changes the model's state — or its actual parameters — while it reads. Instead of holding a million tokens in attention, the system compresses what it has seen into a smaller learned state and updates that state as new tokens arrive. The window stays bounded; the "memory" lives in weights that move during inference.

That third family is the one people mean when they say "the model learns while you use it," and it's where the two papers in your reading list live.

Test-time training: turn the hidden state into a tiny model

The cleanest way to understand test-time training (TTT) is through sequence models generally. A Transformer remembers the past by keeping every past token around (the KV cache) and attending over all of them: perfect recall, linearly growing memory, quadratic compute. An RNN remembers the past by squeezing it into a fixed-size hidden state: cheap and constant-memory, but a fixed vector is a lossy bottleneck.

Sun et al. ("Learning to (Learn at Test Time)," 2024) make a sharp move: what if the hidden state is itself a small model, and the update rule that folds in each new token is a step of gradient descent? Their TTT layers replace the RNN's hidden vector with the weights of a tiny network. As the sequence streams in, each token produces a self-supervised loss, and the layer takes a gradient step on its own inner weights. The past isn't stored token-by-token; it's compressed into the parameters of a model that keeps learning as it reads.

The payoff is asymptotics. Attention's cost per token rises with how much you've already read; a TTT layer's cost per token is flat, because the state is fixed-size no matter how long the sequence gets. In pseudocode the loop is almost embarrassingly simple:

# W is the "hidden state" — the inner model's weights
for x in sequence:                 # streaming tokens
    loss = self_supervised(W, x)   # e.g. reconstruct a corrupted view of x
    W = W - lr * grad(loss, W)     # one SGD step *during inference*
    y = predict(W, x)              # emit output using the just-updated state

Nothing here is trained offline in the usual sense. The inner learning happens at test time, on your specific sequence, which is where the name comes from.

Titans: a long-term memory that decides what's worth keeping

Behrouz, Zhong, and Mirrokni ("Titans: Learning to Memorize at Test Time," 2024) push the idea into a full architecture with an explicit long-term memory module: a neural network that learns to memorize during use and is queried like a lookup. Two of their design choices are worth stealing as mental models even if you never touch the code.

First, surprise as a write signal. A stream of boring, predictable tokens shouldn't overwrite your memory; a surprising one should. Titans measures surprise by the gradient of the memory's loss. A token the memory predicts poorly produces a large gradient, and that's exactly the token worth storing. This is a principled answer to "what do I remember?" that beats "remember the last N tokens."

Second, forgetting is a feature, not a bug. A memory that only ever writes fills up and turns to mush. Titans includes a decay/forget mechanism so old, unreinforced information fades, which lets the same fixed-size memory operate over very long streams without saturating. The paper reports scaling to context windows past two million tokens while keeping the memory's footprint bounded, and it splits the labor three ways: short-term attention for local precision, the long-term neural memory for the distant past, and a persistent component for task knowledge that shouldn't drift.

The honest caveat: these are recent research systems (2024–2025), not the default inside the model you call over an API today. Treat them as where the field is heading, not as a library you drop in this afternoon.

The limits that survive an infinite stream

Here is the part the headline hides. Even a theoretically endless input does not buy you an all-knowing model, because six constraints don't disappear when the window grows.

  • Capacity. A fixed-size state — the whole point of TTT and Titans — can hold only so much before new information overwrites old. You are choosing a lossy compression, not escaping compression.
  • Retrieval accuracy. Long-context models famously degrade in the middle. The "needle in a haystack" tests, and later the NoLiMa benchmark, show recall dropping as relevant facts move away from the edges of a long prompt. More context is not more usable context.
  • Forgetting. The same decay that keeps memory healthy will eventually drop something you needed. Whether that's graceful or catastrophic depends entirely on your write and forget policies.
  • Latency. Test-time gradient steps and million-token attention both cost wall-clock time per token. A memory that updates as it reads is doing extra math on the critical path of every response.
  • Compute and dollars. Long context is priced per input token. Adaptation is priced in FLOPs. Neither is free, and both scale with exactly the thing you were hoping to make unlimited.
  • Security. Any memory that writes what it reads is a persistence layer for whatever it reads. Prompt injection stops being a one-shot trick and becomes a stored payload: poison the memory once, and it can steer later, unrelated sessions. Memory turns a transient attack into a durable one.

What to actually build on Monday

Reach for external memory (RAG) first. It's the option where you control capacity (your database), you can measure retrieval accuracy directly (did the right chunk come back?), forgetting is an explicit delete, and — the part that matters when something goes wrong — the memory is inspectable and scrubbable. You get most of the practical value of "the model knows my documents" without betting on research-grade architectures.

Reach for long context when the task genuinely needs global reasoning over one coherent artifact — a whole codebase, a long contract — where chunking would sever dependencies that retrieval can't reassemble. Pay the quadratic cost knowingly, and test recall in the middle of your real prompts, not just the ends.

Watch test-time adaptation (TTT, Titans) as the direction that dissolves the window-versus-state tradeoff, and read the papers so you're not fooled by the next "infinite memory" pitch. When someone sells you unlimited recall, ask the six questions above. If they can't tell you where the capacity, the forgetting policy, and the injection surface live, they're selling you the slide, not the system.

Sources