The beginner series · 6 min

RAG vs. Embeddings vs. Vectors: WTF Is the Difference?

How AI can search external knowledge.

They are related. They are not competitors. A vector is just a list of numbers. An embedding is a vector that captures the meaning of something — text, usually — as numbers you can compare. RAG is the pattern of finding relevant text and handing it to the model before it answers; it often uses embeddings to do the finding, but it doesn't have to. So they stack rather than compete: RAG can use embeddings, and embeddings are stored as vectors. Asking which one you should “use” is like asking whether you need the library, the catalog, or the shelves.

Reading promise: No assumed AI knowledge. Jargon gets translated before it gets used.

The 30-Second Version

If you hear someone say, “We embedded our documents into a vector database and added RAG,” translate it like this:

DOCUMENTS
   │
   ├─ split into useful chunks
   ▼
EMBEDDING MODEL
   │
   ▼
VECTORS  [0.18, -0.42, 0.91, ...]
   │
   ▼
VECTOR STORE / SEARCH INDEX

User question
   │
   ├─ embed the question too
   ▼
Find similar chunks
   │
   ▼
Give retrieved text to the LLM
   │
   ▼
Generate a grounded answer

The retrieve → add context → generate pattern is RAG.
TermPlain English
VectorA list of numbers.
EmbeddingA useful representation of something encoded as a vector.
Vector searchFind vectors that are close to another vector according to a distance/similarity measure.
Vector databaseInfrastructure designed to store, index, filter, and search vectors.
RAGRetrieve relevant information and give it to a generative model at answer time.

Start With the Least Magical Thing: a Vector

A vector is not an AI brain. It is a list of numbers. In two dimensions, [2.1, 7.4] could be a point on a map. In an embedding model, a vector may have hundreds or thousands of dimensions.

Pretend meaning can be drawn on a 2-D map:

       animals
          ↑
   dog •  • puppy
          │
──────────┼────────────→ machines
          │       • car
          │          • truck
          │
    • chocolate-cookie

Real embeddings do not come with a neat axis labeled “animals.” The model learns a high-dimensional representation. What matters is the relationship between vectors: items with related content often wind up closer together.

So What Is an Embedding?

An embedding model turns an input—commonly text, but sometimes images or other data—into a numeric vector designed to preserve useful relationships. OpenAI’s embedding documentation describes embeddings as numerical representations that can be used to measure relatedness and power search, clustering, recommendations, anomaly detection, and classification.[1]

"Golden retrievers are friendly dogs."
            │
            ▼
       embedding model
            │
            ▼
[0.0218, -0.0047, 0.0381, 0.0924, ...]

Do not read one number and imagine it means “dog.” The information is distributed across the vector. The useful property is that semantic relationships can be compared mathematically.

Keyword search is excellent and still matters. But it primarily rewards lexical overlap. Semantic search tries to retrieve text that is related in meaning even when the words differ.

QUERY
"When did humans first land on the moon?"

DOCUMENT
"The first lunar landing occurred in July 1969."

Few exact words match. The meanings do.

OpenAI’s retrieval documentation describes vector stores as indices for data that can support semantic search; modern retrieval systems can also combine semantic and keyword techniques.[2]

What a Vector Database Actually Does

If you have millions of embedded chunks, you need efficient storage and nearest-neighbor search. A vector database—or a conventional database with vector capabilities—indexes those vectors so you can ask, “Which stored items are closest to this query vector?”

QUERY VECTOR
[0.18, 0.90, -0.41, ...]
       │
       ▼
VECTOR INDEX
       │
       ├─ Handbook §8.4     very similar
       ├─ Benefits FAQ      similar
       ├─ Vacation policy   somewhat similar
       └─ Printer manual    not similar

For example, pgvector adds vector similarity search to PostgreSQL and supports exact and approximate nearest-neighbor search.[3] The database is not “thinking.” It is doing retrieval.

Now RAG Finally Makes Sense

RAG stands for Retrieval-Augmented Generation. The foundational 2020 RAG paper combined a generative model with information retrieved from an external dense index.[4] In beginner English: before asking the model to answer, find useful evidence and put that evidence in front of it.

WITHOUT RAG
Question → model's existing context/learned behavior → answer

WITH RAG
Question → retrieve relevant evidence → put evidence in context → answer

An Employee-Handbook Example

Suppose a company handbook says:

Eligible employees receive sixteen weeks of paid parental leave following the birth, adoption, or placement of a child.

A user asks, “How long can I take off after having a baby?” The wording does not match exactly, but semantic retrieval can surface the parental-leave chunk. The application then sends the question plus that chunk to the LLM.

SYSTEM: Answer using the supplied policy.

CONTEXT: Eligible employees receive sixteen weeks...

USER: How long can I take off after having a baby?

LLM: Eligible employees receive sixteen weeks of paid parental leave.

The Two Phases of a Typical RAG System

Phase 1: Prepare the knowledge

1. Load the documents or records.

2. Extract or normalize the useful content.

3. Split large content into chunks.

4. Create an embedding for each chunk.

5. Store the vector alongside the original text and metadata.

{
  text: "Paid parental leave is sixteen weeks...",
  vector: [-0.31, 0.62, 0.51, ...],
  source: "Employee Handbook",
  page: 47,
  version: "2026-07"
}

Phase 2: Answer a question

1. Create a representation of the user’s query.

2. Retrieve likely-relevant chunks.

3. Optionally filter/rerank them.

4. Put the best evidence in the model’s context.

5. Ask the model to answer from that evidence.

Important: RAG Does Not Require Vectors

RAG requires retrieval, not necessarily embeddings. Retrieval might use keyword search, SQL, metadata filters, APIs, web search, graph queries, vector search, or a hybrid of several methods.

RAG
 └─ retrieval
     ├─ keyword search
     ├─ SQL / metadata filters
     ├─ API lookup
     ├─ graph lookup
     └─ semantic vector search
          └─ embeddings
               └─ vectors
Saying “RAG vs. embeddings vs. vectors” is a little like saying “library research vs. catalog numbers vs. shelves.” One can contain or use the others.

Does the Vector Contain the Original Text?

Not like reversible encryption. Applications commonly store the source text alongside its vector. The vector helps locate information; the original chunk provides the words the model should actually read.

VECTOR → helps find WHERE to look
TEXT   → tells the model WHAT was said

Does RAG Eliminate Hallucinations?

No. It can improve grounding, but the retrieval system can return the wrong document, an outdated policy, a poisoned document, or two contradictory sources. The model can also misread good evidence. A production RAG system needs document governance, chunking, filters, ranking, permissions, citations, and evaluation—not merely “a vector database.”

A Natural Cybersecurity Example

Imagine an internal investigation assistant with incident reports, runbooks, detection documentation, vendor knowledge, and threat intelligence. An analyst asks:

Have we previously seen Microsoft Word launch PowerShell followed by an outbound connection to Dropbox?
Question
Search prior incidents + runbooks + detections
Retrieve the handful of relevant records
LLM compares patterns and summarizes evidence
Answer cites the underlying incidents

That is why RAG became so useful for enterprise AI: it lets a model reason over information that is private, large, or changing without pretending that all of that material permanently lives inside the model.

The Cheat Sheet

ConceptRemember it as
VectorNumbers
EmbeddingMeaning/data → numbers
Vector searchFind nearby numbers
Vector databaseStore and search lots of those numbers
RAGFind useful evidence → show the model → generate an answer

Sources

RAG vs. Embeddings vs. Vectors: WTF Is the Difference? — WTF Is…? — Plain-English AI · AdversariaLLM