Foundational Mechanics · 7 min
Embeddings: turning tokens into geometry
How language models turn discrete tokens into dense vectors where distance means similarity — and where that geometry quietly lies to you.
A model can't multiply words. At the very bottom of every language model, the symbol cat has to turn into a list of floating-point numbers, because matrices only speak numbers. That conversion is an embedding. The trick is that the numbers aren't arbitrary. They're chosen so geometry carries meaning: words used in similar ways land near each other, and the directions between points line up with relationships you'd recognize. This is where "language" first becomes "math," and getting the intuition right pays off everywhere downstream, in attention, retrieval, and classification alike.
The mental model: a learned map
Start with the dumb baseline. Say your vocabulary has 50,000 tokens. The obvious encoding is one-hot: cat is a 50,000-long vector that's 1 in slot 8,213 and 0 everywhere else. It's honest and useless. Every word sits exactly the same distance from every other word. cat and dog are as unrelated as cat and bureaucracy, because any two distinct one-hot vectors are orthogonal and their dot product is 0. You've encoded identity and thrown away all similarity.
An embedding replaces that with a dense vector, say 300 numbers, all nonzero, living in a continuous space the model learned. Nobody hand-assigns the axes. During training the model nudges each word's coordinates so words that show up in similar contexts drift together. What you get is a map where proximity means something.
That dense vector isn't consumed once and thrown away, either. In the next lesson you'll see it's the initial value of the residual stream — the single running per-token vector every transformer block reads from and writes back to — so the embedding is literally where each token's computation begins, not a separate object the network forgets after layer one.
Two things named embeddings: The vector this section is about — the model's internal token embedding, a token id mapped to the representation the Transformer consumes — is a related but different object from the sentence or document embedding a vector database stores for retrieval. They share the core idea (text becomes geometry) and sometimes even a training lineage, but they're produced by different models trained for different jobs and live in different spaces. Don't assume the numbers inside the model are the ones sitting in your vector DB.
one-hot (sparse, orthogonal) embedding (dense, learned) cat = [0 0 1 0 0 0 ... 0] cat = [ 0.21 -0.44 0.90 ... ] (300 dims) dog = [0 0 0 0 1 0 ... 0] dog = [ 0.19 -0.40 0.85 ... ] car = [0 1 0 0 0 0 ... 0] car = [-0.71 0.33 0.10 ... ] all pairwise distances equal cat≈dog (close), cat≠car (far)
The classic demonstration is word2vec, from Mikolov and colleagues at Google in 2013. The insight was almost embarrassingly cheap: predict a word from its neighbors (or the neighbors from the word), do it over billions of words, and the byproduct is a set of vectors with startling structure. Their paper reports learning high-quality vectors from a 1.6-billion-word dataset in under a day, which turned heads in 2013. The party trick is vector arithmetic. Take the vector for king, subtract man, add woman, and the nearest point is queen. Relationships became directions you could add and subtract: a "royalty" direction, a "plural" direction, a "capital-city" direction.
The mechanics: how distance gets measured
Once words are points, you need a ruler. In practice the ruler is usually cosine similarity, the cosine of the angle between two vectors — though it is not the only choice: a raw dot product (identical to cosine once vectors are length-normalized) and plain Euclidean distance are both common, and which one an index uses is a real configuration decision. It ignores length and measures only direction.
cos(a, b) = (a · b) / (|a| · |b|) range: -1 (opposite) to 1 (identical)
Work a 2-D example by hand so the number stops being abstract. Say the model learned:
cat = [0.9, 0.1] dog = [0.8, 0.2] car = [0.1, 0.9]
For cat vs dog: dot product = 0.9·0.8 + 0.1·0.2 = 0.74. The magnitudes are √0.82 ≈ 0.906 and √0.68 ≈ 0.825, so cosine = 0.74 / 0.747 ≈ 0.99. For cat vs car: dot = 0.9·0.1 + 0.1·0.9 = 0.18, magnitudes both ≈ 0.906, cosine = 0.18 / 0.82 ≈ 0.22. The geometry did its job. Near-parallel for the two animals, nearly a right angle for the vehicle.
Why cosine and not plain Euclidean distance? Because in these spaces direction tends to carry the semantics while magnitude often tracks incidental things like word frequency. Normalize the length away and you compare meaning, not loudness. One practical payoff: if you L2-normalize every vector to unit length up front, cosine similarity collapses into a plain dot product. That's exactly why vector databases store normalized vectors and run raw dot products at query time.
Builder: A raw cosine number means nothing in isolation. It's only interpretable relative to the distribution of your own data. In many embedding spaces even unrelated sentences score 0.3 to 0.5, and genuinely related ones sit at 0.7 and up. Calibrate against your corpus before you hard-code a threshold. A "0.8 means duplicate" rule copied from a blog post will quietly betray you.
Sentences, and what a modern model hands you
Word2vec gives one fixed vector per word, a static embedding, and that's its ceiling. The word bank has exactly one vector, so "river bank" and "bank account" collapse onto the same point. Modern embeddings are contextual. They run the whole sentence through a transformer, so bank comes out at different coordinates depending on its neighbors. Ethayarajh's 2019 analysis measured how strong that effect is: in the upper layers of models like BERT and GPT-2, a word's representations across different sentences grow far less similar to each other than they are lower down, though a word still resembles itself more than it resembles other words.
When you call a sentence-embedding model, from the sentence-transformers family descended from Sentence-BERT (Reimers and Gurevych, 2019), you hand it a string and get back one fixed-length vector for the whole thing. all-MiniLM-L6-v2 returns 384 dimensions; all-mpnet-base-v2 returns 768. Two paraphrases land close, two unrelated sentences land far. That's the entire product: text in, one vector out, cosine-comparable.
The gap people trip on: embedding is not retrieval
Here's the distinction that separates people who ship working search from people who file confused bug reports. An embedding is a representation. Retrieval is a system built on top of it. The model gives you a vector. It does not give you an index, a similarity threshold, a chunking strategy, a way to handle "the answer spans three documents," or any notion of why two things are close. Cosine similarity measures relatedness, not relevance. The antonyms "hot" and "cold" are highly related and sit close together, which is great for a thesaurus and a disaster if your user asked for hot drinks and got iced ones. This distance-means-similarity geometry is exactly the machinery RAG runs on, which is why the context-window lesson circles back to it when it explains why RAG exists.
Defender: That "meaning as geometry" property is an attack surface. Because retrieval ranks by cosine proximity, an attacker who controls indexable text can craft a passage engineered to sit near likely queries, get pulled into a RAG context window, and poison the model's input without ever touching the model. Treat every embedded document as untrusted and the retriever as part of your trust boundary, not a neutral utility. The RAG lesson picks this up in depth.
Researcher: Cosine is a convenient default, not ground truth. These spaces are anisotropic, meaning vectors cluster into a narrow cone rather than spreading over the sphere, which inflates baseline similarities and makes raw scores hard to read. If you're benchmarking, whiten or mean-center your embeddings first, and report the distribution of scores rather than a single number.
A five-minute lab
Install sentence-transformers, encode three strings, and print the cosine matrix:
"How do I reset my password?" A "I forgot my login credentials" B "What time does the store close?" C expected: cos(A,B) high (~0.6–0.7), cos(A,C) and cos(B,C) low (~0.1)
Then break it on purpose. Encode "not good" and "good" and watch them score far higher than the meanings deserve. Short negations are a known weak spot, because the surrounding words dominate the vector. That failure is the whole lesson in miniature: the geometry is powerful, but it's approximate, and knowing where it lies to you is the difference between a demo and a system.
Sources
- Mikolov, Chen, Corrado, Dean (2013), Efficient Estimation of Word Representations in Vector Space, arXiv:1301.3781.
- Mikolov, Sutskever, Chen, Corrado, Dean (2013), Distributed Representations of Words and Phrases and their Compositionality, NeurIPS. — https://arxiv.org/abs/1310.4546
- Mikolov, Yih, Zweig (2013), Linguistic Regularities in Continuous Space Word Representations, NAACL-HLT (the king − man + woman ≈ queen result). — https://aclanthology.org/N13-1090/
- Reimers and Gurevych (2019), Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks, EMNLP. — https://aclanthology.org/D19-1410/
- Ethayarajh (2019), How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings, EMNLP, arXiv:1909.00512.
- SBERT.net documentation and sentence-transformers model cards (
all-MiniLM-L6-v2,all-mpnet-base-v2). — https://www.sbert.net/