The beginner series · 7 min
WTF Is an LLM Actually Doing?
Next-token prediction, attention, and why prediction can look like reasoning.
“It predicts the next token” is true — and so incomplete it usually causes more confusion than it clears up. Here's the fuller version: the model reads your text as tokens, and for each next slot it produces a probability across its entire vocabulary, picks one, appends it, and repeats. To pick well it has to have learned grammar, facts, cause and effect, and your instructions — which is why prediction one token at a time can still look like reasoning. It is not looking up stored sentences; the knowledge is spread across billions of numbers, not filed in a table.
Reading promise: No assumed AI knowledge. Jargon gets translated before it gets used.
The Annoyingly Short Answer
A generative language model repeatedly predicts a probability distribution over what token should come next, chooses a token according to its decoding process, appends it, and repeats.
OpenAI educational material describes GPT-style models as learning from large amounts of text through token prediction, while the original Transformer paper introduced the attention-based architecture underlying the modern family of Transformer language models.[1][2]
“Next-Token Prediction” Sounds Way Dumber Than It Is
Suppose I ask you to finish:
The burglar broke the window because the door was _____.
To predict a plausible completion, you implicitly use grammar, the physical world, human motives, and the previous sentence. Maybe “locked.” Next-token prediction can require a surprising amount of structure about the text and world that produced it.
To predict the next token well, a model may need signals about: - syntax - semantics - style - entities and relationships - code structure - cause and effect - common facts - the user's instructions - what appeared 8,000 tokens earlier
A Language Model Is a Probability Machine
Very roughly, the model receives a sequence of tokens and computes scores for possible next tokens. A softmax-like transformation turns scores into a probability distribution. Generation samples or selects from that distribution according to decoding settings.
CONTEXT: "PowerShell launched from WINWORD.EXE is" possible continuation rough intuition --------------------------------------------- " suspicious" high " normal" lower " purple" extremely low The actual model considers an entire vocabulary, not three words.
Where Do Those Probabilities Come From?
From a neural network containing learned parameters. Each layer transforms the representation of the token sequence. During training, the parameters were adjusted so the model became better at predicting appropriate tokens across huge numbers of examples. Articles 7 and 8 unpack that machinery.
Tokens Become Vectors First
A token ID is not useful to a neural network by itself. The model maps tokens into learned numeric representations—embeddings—so computations can operate on them.
Yes, this is related to the embeddings from Article 1. The exact representations and use differ, but the core idea is familiar: turn symbolic inputs into vectors that neural networks can manipulate.
Then Attention Lets Tokens Look at Other Tokens
The Transformer’s defining idea is attention. Instead of processing each word in isolation, the network computes relationships between positions so a token representation can incorporate information from other relevant tokens. The 2017 “Attention Is All You Need” paper introduced the Transformer architecture around attention mechanisms.[2]
Sentence: "Alice told Bob that she would reset her password." When processing "she", useful attention may connect it strongly to "Alice" rather than "password" or "Bob". This is conceptual—not a promise that one attention head cleanly means "pronoun resolver."
Self-Attention, in Human Terms
For each token position, the model produces mathematical objects commonly called a query, key, and value. The query asks, in effect, “what information is useful to me?” Keys help determine which other positions are relevant. Values carry the information to mix together. The model computes attention scores and combines information accordingly.
Token A: QUERY ─┐
compare against KEYS of other tokens
Token B: KEY/VALUE│
Token C: KEY/VALUE│ → attention weights → weighted information mix
Token D: KEY/VALUE┘Do not interpret query/key/value as English questions and answers. They are learned vectors participating in matrix operations. The analogy only gives intuition.
One Attention Pass Is Not the Whole Model
Modern Transformer models stack many blocks. A simplified block contains attention plus feed-forward neural-network operations, normalization, and residual connections. Representations become progressively transformed across the stack.
Why Does the Model Generate One Token at a Time?
During autoregressive text generation, each new token becomes part of the context for the next prediction. That means a response is built sequentially.
This is one reason output length affects latency: text generation must produce a sequence over time rather than calculating an entire long answer in one single choice.
Temperature: Why the Same Prompt Can Produce Different Answers
If several next tokens are plausible, decoding can be more conservative or more exploratory. Temperature is one common control: lower values concentrate probability mass toward high-probability choices; higher values flatten the distribution and make less-likely options more reachable. Exact behavior depends on the API/model.
LOWER RANDOMNESS "The sky is" → " blue" nearly every time HIGHER RANDOMNESS "Write a strange opening line" → more varied plausible choices
So Is It “Just Autocomplete”?
Calling an LLM “autocomplete” is technically suggestive but psychologically misleading. Your phone keyboard predicts next words using a much smaller model and much less context. A modern LLM has learned enormous, layered statistical structure and can condition on long instructions, examples, tools, images, and retrieved evidence.
“It predicts the next token” describes the objective/interface. It does not tell you how much computation or learned structure is required to make a good prediction.
Does the Model Have a Database of Sentences Inside It?
Not in the ordinary database sense. Parameters encode learned statistical relationships distributed across the network. Models can sometimes reproduce memorized material, but it is wrong to imagine a hidden table where row 7,392 contains “Paris is the capital of France.”
BAD MENTAL MODEL question → lookup exact stored sentence → print sentence BETTER MENTAL MODEL question tokens → network computation using learned parameters → next-token probability distribution → generate sequence
What About Reasoning?
This is where language gets contentious. From the outside, models can perform multi-step reasoning tasks, derive intermediate results, use tools, and correct themselves. Mechanistically, those behaviors are still implemented through neural-network computation and token generation. You do not need to choose between “it predicts tokens” and “it can solve reasoning problems.” Both can be true at different levels of description.
A Cybersecurity Example
Give a model:
WINWORD.EXE → powershell.exe -EncodedCommand … → rundll32.exe → outbound connection to a newly registered domain
To produce a good explanation, the model must condition on relationships among process ancestry, command-line behavior, common attacker tradecraft, and the requested format. It is still producing tokens one at a time—but useful prediction depends on learned patterns spanning the entire input.
The Cheat Sheet
| Question | Beginner answer |
|---|---|
| What does an LLM generate? | A sequence of tokens. |
| How? | Repeatedly predicts a distribution over the next token. |
| What gives it capability? | Learned parameters plus neural-network computation over the supplied context. |
| What is attention doing? | Letting token representations incorporate information from other relevant positions. |
| Is it a sentence database? | No. Learned information is distributed through model parameters. |
| Is “just autocomplete” accurate? | Mechanistically adjacent, but far too reductive to explain modern capability. |
Sources
- [1] OpenAI Cookbook — Summarizing long documents (LLM/token prediction overview). — https://developers.openai.com/cookbook/examples/summarizing_long_documents
- [2] Vaswani et al. — Attention Is All You Need. — https://arxiv.org/abs/1706.03762
- [3] Google Machine Learning Crash Course — Introduction to Large Language Models. — https://developers.google.com/machine-learning/resources/intro-llms
- [4] OpenAI Cookbook — Using log probabilities. — https://developers.openai.com/cookbook/examples/using_logprobs