Foundational Mechanics · 7 min

How tokenization works (and why it bites you)

Your text is compiled to integers by a frozen, learned lookup — and every surprise in cost, context, and safety starts at that boundary.

A language model never sees your text. Before a single matrix multiply happens, your string is chopped into pieces called tokens, and each piece is swapped for an integer index into a fixed vocabulary. The model consumes integers, predicts a probability distribution over the next integer, and something downstream turns integers back into characters. Tokenization is the boundary layer between human strings and the model's numeric world, and like most boundary layers, it's where the leaks happen.

If you take one idea from this lesson, take this: a token is neither a word nor a character. It's whatever fragment the tokenizer's training procedure decided was worth its own slot. Sometimes that's a whole common word ( the), sometimes a fragment (token, ization), sometimes a single byte. The rules are learned from data, frozen at training time, and they leak into everything you build on top.

Why subwords at all

The two obvious designs both fail. A word-level vocabulary can't represent anything it didn't see in training — every novel name, typo, or URL collapses to a single "unknown" token, and the vocabulary explodes once you add more languages. A pure character vocabulary never hits an unknown but makes sequences painfully long and forces the model to relearn spelling for every word. Subwords are the compromise: common words stay whole, rare words fracture into reusable pieces, and nothing is ever truly out-of-vocabulary.

Two families dominate, and both trace to specific papers worth knowing.

Byte-Pair Encoding (BPE), adapted for NLP by Sennrich, Haddow and Birch (2016), is a bottom-up merge process. Start with the text as individual symbols (characters, or in modern variants, raw bytes). Count every adjacent pair, merge the most frequent pair into a new symbol, and repeat for a fixed number of merges. That merge count is your vocabulary size.

Corpus: "lower" is common; "newest" and "widest" share "est"

start:   l o w e r   n e w e s t   w i d e s t
merge 1: pair (e,s) most frequent -> "es"
         l o w e r   n e w es t    w i d es t
merge 2: pair (es,t) -> "est"
         l o w e r   n e w est     w i d est
merge 3: pair (l,o) -> "lo"      ...and so on

The learned merge list is the tokenizer. At inference you replay those merges greedily over the input, which is why BPE is deterministic. It's also why a leading space matters: hello and hello sit in different word slots, so the same letters can split differently depending on what precedes them.

Unigram / SentencePiece. Kudo and Richardson (2018) shipped SentencePiece, which trains directly from raw sentences with no pre-tokenization step. That detail is subtle but load-bearing. Earlier BPE tooling assumed the text was already split into words on whitespace, which bakes in an assumption that breaks for Chinese, Japanese, or Thai, where words aren't space-delimited. SentencePiece treats the input as a raw Unicode stream and encodes the space itself as a visible symbol, ▁ (U+2581). So ▁Hello▁world reverses cleanly back to the original bytes, spaces and all. Detokenizing is just concatenation plus replacing ▁ with a space.

SentencePiece can run BPE, but its signature algorithm is the unigram language model from Kudo (2018). Instead of building up from characters, it starts with a large candidate vocabulary and prunes downward: fit a unigram probability to each piece with EM, then repeatedly drop the lowest-probability pieces until you hit the target size. The payoff is that a unigram model can score multiple valid segmentations of a string and pick the most probable. That also enables subword regularization: sampling different segmentations during training to make models more robust to how a word gets split.

Researcher: BPE gives you one deterministic split; unigram gives you a distribution over splits. That gap is exploitable. Subword regularization and BPE-dropout both trade determinism for robustness, and probing a model with alternate segmentations of the same string is a real evaluation technique.

The numbers you actually care about

Vocabulary size is a design knob, not a constant. GPT-2's BPE vocabulary is 50,257 tokens. The cl100k_base tokenizer behind GPT-3.5/4 is about 100k; o200k_base roughly doubles that again. Llama 1 and 2 use a 32,000-piece SentencePiece vocabulary, but Llama 3 jumped to a ~128k tiktoken-style BPE — a reminder that "the Llama tokenizer" isn't one thing. Bigger vocab means fewer tokens per sentence (cheaper, shorter sequences) at the cost of a fatter embedding table and softmax.

For English prose, OpenAI's durable rule of thumb is ~4 characters per token, roughly 75 words per 100 tokens. Do not trust that number anywhere else. It collapses the moment your input stops looking like the training corpus, and that's exactly where the pain starts.

Where it bites

Code and structured text. Whitespace, indentation, and punctuation each cost tokens. A snake_case identifier like get_user_token might split as get, _user, _token, or worse. Pretty-printed JSON burns real tokens on indentation you don't care about, and minifying can measurably cut cost.

URLs and hashes. A random-looking string has no frequent pairs to merge, so it shatters into near-per-character tokens. A UUID or a long signed URL can eat 40–60 tokens for something that's 36 characters on screen.

Non-English languages. Vocabularies are trained on corpora that skew heavily English. The same sentence in English, Hindi, and Burmese can differ by 3–5x in token count. Users in under-represented languages pay more per sentence and get less usable context, a fairness-and-cost problem baked straight into the tokenizer.

Counting and arithmetic. 1234567 might be one token or several depending on how the tokenizer handles digits. Ask a model to reverse a string or count letters and it struggles partly because it can't see letters, only chunks. The "how many r's in strawberry" failure is a tokenization artifact as much as a reasoning one.

Builder: Never size a context budget by counting words or characters. Call the real tokenizer for the exact model you deploy (tiktoken, the HF tokenizers library, or sentencepiece) and count tokens. A prompt that fits your English test can overflow the window on production Japanese input and silently truncate your system prompt.

The security angle

Token boundaries are an attack surface, and it shows up three ways.

First, glitch tokens. Rumbelow and Watkins (2023) found strings like SolidGoldMagikarp that had their own single token in GPT-2/3's vocabulary — scraped from Reddit usernames and log dumps — yet almost never appeared in the actual training text. The embedding for such a token is essentially untrained noise, and feeding it makes the model hallucinate, evade, or emit garbage. A vocabulary entry the model never learned is a live blind spot.

Second, boundary framing in prompt injection. Filters pattern-match on strings, but the model reasons over tokens. An attacker can split, pad, or Unicode-decorate a banned phrase so it survives your string-level filter yet tokenizes into pieces the model reassembles into the original meaning. Homoglyphs and zero-width characters shift token boundaries in ways a naive if "ignore previous" in text check never sees.

Third, safety behavior doesn't transfer across tokenizers. Because tokenization changes between model versions, a string that fragments harmlessly under one tokenizer can map to a single evocative token under another. SolidGoldMagikarp fragments into several tokens in newer vocabularies and was one token in the old — same text, different model-internal object.

Filter sees:   "i g n o r e"   (fails substring match "ignore")
Model sees:    [ign][ore]       -> reassembles meaning -> obeys
Defender: Run content and injection checks on tokenized input, or normalize hard (NFKC, strip zero-width, collapse whitespace) before both your filter and the model see the text. Match what the model matches. And when you swap model versions, re-run your red-team corpus, because the tokenizer changed underneath you.

We'll return to this in the prompt-injection module, where token boundaries stop being a curiosity and become the primary evasion primitive. For now, keep the mental model: your text is compiled to integers by a frozen, learned lookup, and every surprise in cost, context, and safety starts there.

Sources