Start Here — LLMs in Plain Language · 5 min
How a model runs: from your message to its reply
Follow a single message from the moment you hit send to the reply that streams back, and you'll understand what a language model is actually doing the whole time.
You type a message, hit send, and words appear one at a time, as if something on the other end is thinking them up on the spot. It looks like magic, or like a person typing back. It's neither. It's a specific, repeatable process, and once you can picture it, most of the strange behavior you've seen from these tools stops being strange.
Let's follow one message all the way through.
If you take one idea from this lesson, take this: the model builds its reply one small piece at a time, and each new piece is chosen by looking at everything that came before it. That single loop explains why replies stream in, why longer conversations get slower, and why the thing forgets you the moment you open a new chat.
First, your words become numbers
A language model doesn't read text the way you do. It can only work with numbers. So before anything else happens, your message gets chopped into small chunks, and each chunk is swapped for a number.
Those chunks are called tokens — a token is a piece of text, often a whole common word but sometimes just a fragment like ing, or a single character. "Cats" might be one token; an unusual name might get split into three or four. You don't need the details yet; the next module is entirely about this step, because it quietly shapes more than you'd expect. For now, just hold the picture: your sentence goes in as text and comes out as a list of numbers the model can actually process.
Then it predicts the next piece, over and over
Here's the engine at the center of the whole thing. A language model does exactly one job: given everything so far, work out what comes next.
That's it. It looks at the numbers standing in for your message and rates, for every possible next token, how likely each one is to follow. Then it picks one of the likely ones — usually with a small dose of deliberate randomness mixed in, which is why you can ask the exact same question twice and get two different answers. The chosen token gets added to the end of the text, and the model runs the whole thing again to choose the piece after it. And again. And again.
Picture it like this. Imagine someone who has read an enormous amount of text and is very good at the game "what word comes next." You hand them the start of a sentence; they add one word. Then they read the whole thing again, now including their own word, and add one more. They keep going, always feeding their latest word back into what they're reading, until the thought is finished.
That feedback step is why replies stream — why you see the answer arrive word by word instead of all at once. You're not watching a loading bar uncover a finished paragraph. You're watching the pieces get made, live, one prediction at a time. The word already on your screen is part of what the model is now reading to choose the next one.
Where the analogy breaks down: the "person reading" isn't looking anything up or checking facts. Each pick is a statistical bet about what tends to come next, learned from patterns in mountains of text. That's the honest reason these models sometimes state wrong things with total confidence — a fluent, likely-sounding next word is not the same as a true one. Keep that in mind whenever an answer actually matters.
This takes a staggering amount of arithmetic
Every one of those next-token bets isn't a quick lookup. It's an enormous pile of multiplication and addition — billions of tiny calculations — run across the model's internal values to weigh what should come next. And that whole pile runs again for every token in the reply.
That's why this doesn't happen on an ordinary laptop chip. It runs on GPUs (graphics processing units) — specialized hardware first built to draw video-game graphics, which happen to be very good at doing huge batches of simple math all at once. The same trait that paints a fast game scene is exactly what a language model needs. GPUs are the reason a long answer arrives in seconds instead of minutes.
Using the model vs. building it: inference and training
What we've just described — feeding in your message and generating a reply — is called inference: using an already-finished model to produce an answer. Every chat you've ever had was inference.
Inference is not where the model learned anything. That was training — a separate, earlier, one-time process where the model was built by having it read through vast amounts of text and slowly adjusting its internal values until it got good at the next-token game. Training a large model takes weeks or months, thousands of those GPUs running together, and enormous cost. Inference reuses that finished result cheaply, over and over. (A later module covers training properly; for now, just don't mix the two up.)
One consequence matters right away: the model's knowledge froze when training ended. It has no live line to the world and can't know anything that happened after that cutoff — unless the app around it goes and fetches fresh information to hand in alongside your message.
Its memory is just the current conversation
So how does it "remember" what you said three messages ago? Everything the model can see for this reply — your latest message plus the earlier back-and-forth — is called the context window. Think of it as a whiteboard the model can read while it works: your conversation is written on it, and that's the model's entire short-term memory for the task.
The whiteboard has an edge. Fill it up and the oldest lines get pushed off the top — one reason very long conversations start losing track of earlier details.
Now the part that trips everyone up: the model itself keeps nothing between messages. It's stateless — it doesn't quietly hold your conversation in its head. A chat feels continuous only because the app resends the whole history — every earlier turn — stacked in front of your new message each time. Start a fresh chat and that history isn't sent, so from the model's side you're a total stranger. It's not being forgetful. There was never anything stored to forget.
The whole loop, in one picture
Put it together, and here's what happens when you hit send:
- Your text is split into tokens and turned into numbers.
- The full conversation so far — the context window — goes to a model running on GPUs.
- The model rates the likely next tokens, picks one, adds it to the end, and repeats — that's the reply appearing word by word.
- It stops when the answer is complete, and remembers none of it.
That loop is the beating heart of every chatbot you'll ever use. The rest of this course is really just zooming in on its parts — starting, next, with those tokens.