RAG & Knowledge Systems · 6 min
Why RAG Exists: Parametric Weights vs. Retrieved Context
Two hard limits — a fixed context window and weights frozen at training cutoff — force a split between knowledge baked into weights and a corpus fetched at inference.
Module 1 left you with two structural facts about a transformer, and both are load-bearing here. First, the context window is fixed. The model attends over a bounded number of tokens and not one more. Second, the weights are frozen at a training cutoff. Whatever the model "knows" got compressed into its parameters during training and is, from then on, a read-only snapshot of a past world.
RAG is the engineering response to exactly those two limits. It isn't a clever trick; it's a direct consequence. Once you accept that you can't retrain the model for every new fact and you can't paste your entire corpus into the prompt, the shape of the solution is nearly forced. Keep the reasoning ability in the weights, move the facts to an external store, and fetch just the relevant slice into the window at query time.
Two kinds of memory
The framing that makes this click comes from Lewis et al. (2020), the paper that named the pattern. It splits a model's knowledge into two memories.
- Parametric memory is knowledge encoded in the weights. It's spread across billions of parameters and accessed by a forward pass, with no lookup step. That makes it fast and always available, but also fundamentally lossy: the model learned statistical regularities, not a verbatim record. Ask it for the population of a specific town or the signature of an API added last quarter and it will often produce something fluent and wrong. That failure is confabulation, and you can't prompt it away. It's what interpolation over compressed training data does at the tails.
- Non-parametric memory is an explicit external corpus: documents, passages, rows. Nothing is baked in. You look things up at inference and drop the results into the context window. Updating it means editing a store, not running a training job.
The distinction earns its keep because the two memories fail in opposite ways. Parametric memory is stale and confident. Non-parametric memory is current but inert; it does nothing until something retrieves from it and hands it to a generator. RAG is the wiring between them.
PARAMETRIC (weights) NON-PARAMETRIC (corpus)
┌───────────────────┐ ┌────────────────────────┐
│ fluency, grammar, │ │ your docs, tickets, │
│ reasoning, "how │ │ policies, code, facts │
│ language works" │ │ as of five minutes ago │
├───────────────────┤ ├────────────────────────┤
│ frozen at cutoff │ │ edit anytime │
│ lossy on specifics│ │ verbatim, citable │
│ no source │ │ has a source URI │
└───────────────────┘ └────────────────────────┘
reasoning ◄────── RAG glues them ──────► factsThe base pattern: retrieve, then generate
Strip RAG to its skeleton and it's two steps.
query ──► [ RETRIEVER ] top-k passages
embed q, search index ──────────► p1 p2 ... pk
│
prompt = q + p1..pk ▼
answer ◄── [ GENERATOR ] ◄──────────────── in-context
(frozen LLM reads the passages, writes the answer)The retriever turns the query into a vector and runs a nearest-neighbor search over a pre-embedded corpus, returning the top k passages. The generator, an ordinary frozen LLM, receives the original query with those passages pasted in front of it and writes an answer conditioned on them. The model's weights never move. The only thing that changed is what's in the window.
Everything else in this module (chunking, embedding models, hybrid search, rerankers, evaluation) refines one of those two boxes. Get the skeleton right and the rest is tuning.
What Lewis actually built
Concreteness helps, so here are the real numbers from the 2020 paper instead of a hand-wave. The non-parametric store was a December 2018 Wikipedia dump, sliced into disjoint 100-word chunks: about 21 million passages, each embedded once into a dense vector index. The retriever was DPR (Dense Passage Retrieval), a BERT bi-encoder that maps query and passage into a shared vector space so relevance becomes a dot product, with the search itself done by maximum-inner-product search over the index. The generator was BART, a seq2seq model. At query time the system pulled the top k passages (the paper reports values like 5 and 10) and generated conditioned on them.
The payoff was measurable. On Natural Questions, RAG-Sequence hit 44.5 exact match. The yardstick the paper picks is T5-large at 28.9, the closed-book model closest in size to RAG's generator, answering from parameters alone. That gap of roughly 16 points isn't about scale, since the parameter counts sit in the same neighborhood; it's about the corpus. RAG also set state of the art on three open-domain QA tasks at the time. The parameters didn't get smarter. The model just stopped guessing at facts it could instead read.
One detail worth carrying forward: Lewis proposed two variants. RAG-Sequence commits to a single retrieved passage for the whole answer. RAG-Token can attend to a different passage token by token. That "which passage for which span" question comes back constantly once answers need to fuse multiple sources.
Researcher: The honest comparison isn't RAG versus no-RAG, it's RAG versus closed-book: a model answering the same questions from parameters only, at comparable size. That isolates what retrieval contributes from what the base model already knew. Keep a closed-book baseline in every eval. It's the denominator that keeps you honest.
Why not just make the window bigger
Fair question, and the usual first instinct. Context windows keep growing, so why retrieve at all, why not stuff the whole corpus in? Three reasons it doesn't scale.
Cost and latency track the number of tokens in the window. A corpus of millions of passages will never fit, and even when a big chunk does fit, you pay for every token on every call. Second, models degrade with irrelevant context: bury the answer in the middle of a huge dump and accuracy drops, the "lost in the middle" effect you'll meet later in this module. Retrieval is a relevance filter that keeps the window dense with signal. Third, and this is the one people underweight, a corpus you retrieve from is a corpus you can update, audit, and cite. Edit a document and the next query sees the change. No retraining, no redeploy.
Builder: This is the operational win. New facts land in the store, not in a fine-tune. The freeze on the weights stops being a wall; you route around it by editing an index. Your "knowledge deploy" becomes a document write.
Defender: Non-parametric memory is also an input path into the prompt. Anything the retriever can pull, an attacker who can write to the corpus can plant. That's indirect prompt injection, and it rides in on the exact channel RAG opens. Treat retrieved passages as untrusted input, not trusted context. Provenance on every chunk (where it came from, who could write it) is a security control, not a UX nicety.
The mental model to keep
Hold onto this for the rest of the module: weights are for skills, the corpus is for facts. Grammar, reasoning, how to structure an answer are parametric. Anything specific, current, private, or citable is non-parametric, fetched at inference. When a RAG system is wrong, the split tells you where to look. A fluent-but-fabricated answer usually means retrieval failed and the model fell back on lossy parameters. A correct-but-garbled one usually means retrieval worked and generation dropped the ball.
Next up: how the retriever actually turns text into vectors, and why the choice of embedding model quietly decides how good every downstream answer can be.
Sources
- Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., Kiela, D. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 33, 9459–9474. https://arxiv.org/abs/2005.11401
- Karpukhin, V., et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering. EMNLP 2020. https://arxiv.org/abs/2004.04906
- Liu, N. F., et al. (2023). Lost in the Middle: How Language Models Use Long Contexts. https://arxiv.org/abs/2307.03172
- Greshake, K., et al. (2023). Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. https://arxiv.org/abs/2302.12173