The beginner series · 6 min
RAG vs. Embeddings vs. Vectors: WTF Is the Difference?
How AI can search external knowledge.
They are related. They are not competitors. A vector is just a list of numbers. An embedding is a vector that captures the meaning of something — text, usually — as numbers you can compare. RAG is the pattern of finding relevant text and handing it to the model before it answers; it often uses embeddings to do the finding, but it doesn't have to. So they stack rather than compete: RAG can use embeddings, and embeddings are stored as vectors. Asking which one you should “use” is like asking whether you need the library, the catalog, or the shelves.
Reading promise: No assumed AI knowledge. Jargon gets translated before it gets used.
The 30-Second Version
If you hear someone say, “We embedded our documents into a vector database and added RAG,” translate it like this:
DOCUMENTS │ ├─ split into useful chunks ▼ EMBEDDING MODEL │ ▼ VECTORS [0.18, -0.42, 0.91, ...] │ ▼ VECTOR STORE / SEARCH INDEX User question │ ├─ embed the question too ▼ Find similar chunks │ ▼ Give retrieved text to the LLM │ ▼ Generate a grounded answer The retrieve → add context → generate pattern is RAG.
| Term | Plain English |
|---|---|
| Vector | A list of numbers. |
| Embedding | A useful representation of something encoded as a vector. |
| Vector search | Find vectors that are close to another vector according to a distance/similarity measure. |
| Vector database | Infrastructure designed to store, index, filter, and search vectors. |
| RAG | Retrieve relevant information and give it to a generative model at answer time. |
Start With the Least Magical Thing: a Vector
A vector is not an AI brain. It is a list of numbers. In two dimensions, [2.1, 7.4] could be a point on a map. In an embedding model, a vector may have hundreds or thousands of dimensions.
Pretend meaning can be drawn on a 2-D map:
animals
↑
dog • • puppy
│
──────────┼────────────→ machines
│ • car
│ • truck
│
• chocolate-cookieReal embeddings do not come with a neat axis labeled “animals.” The model learns a high-dimensional representation. What matters is the relationship between vectors: items with related content often wind up closer together.
So What Is an Embedding?
An embedding model turns an input—commonly text, but sometimes images or other data—into a numeric vector designed to preserve useful relationships. OpenAI’s embedding documentation describes embeddings as numerical representations that can be used to measure relatedness and power search, clustering, recommendations, anomaly detection, and classification.[1]
"Golden retrievers are friendly dogs."
│
▼
embedding model
│
▼
[0.0218, -0.0047, 0.0381, 0.0924, ...]Do not read one number and imagine it means “dog.” The information is distributed across the vector. The useful property is that semantic relationships can be compared mathematically.
Keyword Search vs. Semantic Search
Keyword search is excellent and still matters. But it primarily rewards lexical overlap. Semantic search tries to retrieve text that is related in meaning even when the words differ.
QUERY "When did humans first land on the moon?" DOCUMENT "The first lunar landing occurred in July 1969." Few exact words match. The meanings do.
OpenAI’s retrieval documentation describes vector stores as indices for data that can support semantic search; modern retrieval systems can also combine semantic and keyword techniques.[2]
What a Vector Database Actually Does
If you have millions of embedded chunks, you need efficient storage and nearest-neighbor search. A vector database—or a conventional database with vector capabilities—indexes those vectors so you can ask, “Which stored items are closest to this query vector?”
QUERY VECTOR
[0.18, 0.90, -0.41, ...]
│
▼
VECTOR INDEX
│
├─ Handbook §8.4 very similar
├─ Benefits FAQ similar
├─ Vacation policy somewhat similar
└─ Printer manual not similarFor example, pgvector adds vector similarity search to PostgreSQL and supports exact and approximate nearest-neighbor search.[3] The database is not “thinking.” It is doing retrieval.
Now RAG Finally Makes Sense
RAG stands for Retrieval-Augmented Generation. The foundational 2020 RAG paper combined a generative model with information retrieved from an external dense index.[4] In beginner English: before asking the model to answer, find useful evidence and put that evidence in front of it.
WITHOUT RAG Question → model's existing context/learned behavior → answer WITH RAG Question → retrieve relevant evidence → put evidence in context → answer
An Employee-Handbook Example
Suppose a company handbook says:
Eligible employees receive sixteen weeks of paid parental leave following the birth, adoption, or placement of a child.
A user asks, “How long can I take off after having a baby?” The wording does not match exactly, but semantic retrieval can surface the parental-leave chunk. The application then sends the question plus that chunk to the LLM.
SYSTEM: Answer using the supplied policy. CONTEXT: Eligible employees receive sixteen weeks... USER: How long can I take off after having a baby? LLM: Eligible employees receive sixteen weeks of paid parental leave.
The Two Phases of a Typical RAG System
Phase 1: Prepare the knowledge
1. Load the documents or records.
2. Extract or normalize the useful content.
3. Split large content into chunks.
4. Create an embedding for each chunk.
5. Store the vector alongside the original text and metadata.
{
text: "Paid parental leave is sixteen weeks...",
vector: [-0.31, 0.62, 0.51, ...],
source: "Employee Handbook",
page: 47,
version: "2026-07"
}Phase 2: Answer a question
1. Create a representation of the user’s query.
2. Retrieve likely-relevant chunks.
3. Optionally filter/rerank them.
4. Put the best evidence in the model’s context.
5. Ask the model to answer from that evidence.
Important: RAG Does Not Require Vectors
RAG requires retrieval, not necessarily embeddings. Retrieval might use keyword search, SQL, metadata filters, APIs, web search, graph queries, vector search, or a hybrid of several methods.
RAG
└─ retrieval
├─ keyword search
├─ SQL / metadata filters
├─ API lookup
├─ graph lookup
└─ semantic vector search
└─ embeddings
└─ vectorsSaying “RAG vs. embeddings vs. vectors” is a little like saying “library research vs. catalog numbers vs. shelves.” One can contain or use the others.
Does the Vector Contain the Original Text?
Not like reversible encryption. Applications commonly store the source text alongside its vector. The vector helps locate information; the original chunk provides the words the model should actually read.
VECTOR → helps find WHERE to look TEXT → tells the model WHAT was said
Does RAG Eliminate Hallucinations?
No. It can improve grounding, but the retrieval system can return the wrong document, an outdated policy, a poisoned document, or two contradictory sources. The model can also misread good evidence. A production RAG system needs document governance, chunking, filters, ranking, permissions, citations, and evaluation—not merely “a vector database.”
A Natural Cybersecurity Example
Imagine an internal investigation assistant with incident reports, runbooks, detection documentation, vendor knowledge, and threat intelligence. An analyst asks:
Have we previously seen Microsoft Word launch PowerShell followed by an outbound connection to Dropbox?
That is why RAG became so useful for enterprise AI: it lets a model reason over information that is private, large, or changing without pretending that all of that material permanently lives inside the model.
The Cheat Sheet
| Concept | Remember it as |
|---|---|
| Vector | Numbers |
| Embedding | Meaning/data → numbers |
| Vector search | Find nearby numbers |
| Vector database | Store and search lots of those numbers |
| RAG | Find useful evidence → show the model → generate an answer |
Sources
- [1] OpenAI — Vector embeddings. — https://developers.openai.com/api/docs/guides/embeddings
- [2] OpenAI — Retrieval. — https://developers.openai.com/api/docs/guides/retrieval
- [3] pgvector — Open-source vector similarity search for PostgreSQL. — https://github.com/pgvector/pgvector
- [4] Lewis et al. — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. — https://arxiv.org/abs/2005.11401