Start Here — LLMs in Plain Language · 7 min

The words everyone uses: a plain-language glossary

The handful of words you'll meet everywhere in AI — model, weights, prompt, token, context, inference, training, hallucination, and "reasoning" — explained through one short chat so they connect instead of float.

You've used a chatbot. You typed something, it typed back, it felt a little like magic. This lesson isn't about the magic — it's about the words. There's a small pile of terms that show up in every article, every product page, every argument online about AI, and if you don't have a plain-English handle on them, all of it reads like static.

So here's the whole rest of the course's vocabulary, and most of the internet's, in one place. We'll hang every word on a single tiny conversation. Picture yourself typing this:

You: Suggest a name for my cat. She's orange and grumpy. Assistant: How about Marmalade? Or, if you want to lean into the grumpiness — Sir Reginald Scowls.

That exchange is enough to explain everything. Let's walk the terms in the order they actually come up.

Model

The model is the thing that produced that reply — a very large file of numbers, plus the software that runs it. It isn't a program someone wrote line by line, with a rule for cats and a rule for names. It's a pattern-predictor: given some text, it predicts what text plausibly comes next. That's it. The whole apparatus is a machine for guessing the next chunk of writing, one chunk at a time, and everything else is built on top of that one trick.

The dominant design under the hood is called a Transformer, an approach introduced by researchers in 2017 that turned out to scale astonishingly well. You don't need the internals yet. Just hold onto this: a model is a big pile of numbers that predicts what comes next.

Parameters (a.k.a. weights)

Those numbers have a name: parameters, also called weights. They're the millions or billions of little dials that got tuned when the model was built, and together they encode everything it "knows" — how language fits together, that cats have names, that orange cats are a running joke. When you hear "a 70-billion-parameter model," that's the count of dials. More dials generally means more capacity, but it's not the whole story, and bigger isn't automatically better.

Take one idea from this section: the model's knowledge lives entirely in these frozen numbers. It didn't look anything up to name your cat. It answered from patterns baked into its weights.

Prompt

Your message — "Suggest a name for my cat. She's orange and grumpy." — is the prompt: the text you hand the model to respond to. Prompts can be a one-liner or pages of instructions, examples, and background. The prompt is your entire steering wheel. The model has no other channel into what you want; it only sees the words you (and the app around you) give it.

Token

The model doesn't read letters or whole words. It chops text into tokens — bite-sized pieces, usually a common word or a fragment of one. "Marmalade" might arrive as Marm + alade; "grumpy" might be a single token. As a rough feel, a token is about four characters of English.

Why care? Because tokens are the unit of nearly everything downstream: how much you can send, how much you pay, how fast a reply comes. When people say a model has a limit of "128,000 tokens," they're counting these pieces, not words. (Tokens are important enough that the very next module is entirely about them.)

Context and the context window

Everything the model can "see" at the moment it answers — your prompt, the assistant's earlier replies, any instructions the app quietly adds — is the context. The context window is the ceiling on how much context fits at once, measured in tokens.

Here's the part that trips up almost everyone: the context window is not memory. It's more like a desk the model works at. Whatever's on the desk right now, it can use. Start a brand-new chat and the desk is wiped clean — the model won't remember your grumpy orange cat, because on its own it carries nothing from one conversation to the next. Each turn, the whole visible conversation gets handed over again from scratch. The illusion of memory inside a single chat is just the app re-showing the model the entire thread every time.

If a chatbot ever seems to recognize you across separate chats, that's a memory feature bolted on around the model — the app saved a few notes about you and pastes them back onto the desk. Handy, but it's the app doing the remembering. The model itself still starts every conversation blank.

Inference

Inference is the act of running the model to get an answer — the few seconds between you hitting enter and "Sir Reginald Scowls" appearing. The model reads your tokens and generates the reply one token at a time, each new token shaped by everything before it. That's why text streams out in a little ribbon instead of landing all at once: you're watching inference happen, token by token.

Inference is also where the ongoing cost lives. Training a model happens once, up front; inference happens every single time anyone uses it.

Training and fine-tuning

Training is how the model got made. It's the long, expensive, up-front process of adjusting all those weights by showing the model staggering amounts of text and having it practice predicting the next token, over and over, until the dials settle into useful patterns. This first phase — call it pretraining — produces a model that's fluent but raw: more of a text-continuer than a helpful assistant.

Fine-tuning is a lighter follow-up round of training that shapes that raw model into something well-behaved — nudging the weights so it answers questions, follows instructions, and declines things it shouldn't. The polite, helpful assistant that suggested a cat name is a pretrained model that was then fine-tuned into an assistant.

One consequence to lock in: training happened at a fixed point in the past and then stopped. On its own, the model has no live connection to the world. Ask it about something that happened after its training ended and — unless the app around it can go fetch the answer and paste it in — it simply doesn't know. Its built-in knowledge has a cutoff date.

Hallucination

Sometimes the model states something false with total confidence — a made-up fact, a fake quote, a citation to a paper that doesn't exist. That's a hallucination, and it's not a bug you can fully patch out. Remember what the thing actually does: predict plausible next text. "Plausible" and "true" usually overlap, but not always, and the model has no built-in sense of the difference. A confident tone is not evidence. This is the single most important limit to carry with you — verify anything that matters.

"Reasoning" model

Lately you'll see models sold as reasoning models. The idea: instead of answering immediately, the model spends extra effort at inference time, working through a problem in steps before committing to a reply. For hard questions — math, multi-step logic — that extra deliberation genuinely helps. But the honest framing is just "more computation spent per answer," not the model suddenly thinking the way you do. It's still predicting tokens; it's simply doing more of that work before it shows you the result.

Quick reference

TermPlain meaning
ModelA big file of numbers that predicts the next chunk of text
Parameters / weightsThe tuned numbers where all the model's "knowledge" lives
PromptThe text you give the model to respond to
TokenA bite-sized piece of text (~4 characters); the unit of length and cost
Context / context windowWhat the model can see right now, and the token limit on it — not memory
InferenceRunning the model to generate an answer, one token at a time
TrainingThe one-time process that set the weights by predicting text
Fine-tuningExtra training that turns a raw model into a helpful assistant
HallucinationA confident, made-up falsehood — a built-in risk, not a rare glitch
Reasoning modelA model that spends extra effort per answer before replying

Keep this page open in a tab. You don't have to memorize it — you'll pick these up by seeing them used. That starts in the very next module, where we crack open the humble token and find out why it's neither a word nor a letter.

The words everyone uses: a plain-language glossary — All About LLMs, from AI to Z · AdversariaLLM