Start Here — LLMs in Plain Language · 5 min
What LLMs are good at — and what they are not
LLMs are fluent-language machines: great at reshaping words, unreliable about facts — and knowing the difference is the whole skill.
A large language model (LLM) — the thing behind ChatGPT and tools like it — does one deceptively simple thing: it reads the text so far and predicts what word-piece is likely to come next. Then it does that again, and again, one piece at a time, until it has written a whole answer. That's it. Everything impressive it does, and every way it lets you down, comes out of that single trick.
If you take one idea from this lesson, take this: an LLM is a language machine, not a fact machine. It is astonishingly good at producing text that reads like a fluent, knowledgeable person wrote it. Whether that text is true is a separate question the model isn't really built to answer.
Why it's good at what it's good at
The model learned to predict text by reading an enormous amount of it. To get good at guessing the next word across billions of examples, it had to soak up the patterns of language — grammar, tone, how a recipe is laid out, what a polite email sounds like, how code is usually written. So the things it does best are really the same skill wearing different hats: taking language and reshaping it.
That covers a huge amount of everyday work:
- Drafting and rewriting — turn rough notes into a clean email; make a paragraph shorter, warmer, or more formal.
- Summarizing — boil a long article or thread down to its main points.
- Explaining — re-explain a confusing idea in plainer words, or with an analogy.
- Translating — between languages, or between "expert" and "beginner."
- Coding help — draft a function, or explain what an error message means.
- Brainstorming — spin up twenty names or angles for you to react to.
- Extracting structure — pull the dates, names, and amounts out of a messy block of text into a tidy list.
- Answering questions — when you hand it the source. Paste in a document and ask about that, and it's on solid ground.
Notice the pattern. In every one of these, the raw material is already in front of the model, and its job is to transform it. That plays straight to its strength. Think of it as an extraordinarily well-read assistant who can rephrase anything you hand them, instantly, in any style. Where that picture breaks: a good human assistant knows when they don't know. This one usually doesn't.
Where it lets you down
Because the model is always just predicting plausible next words, it will confidently produce text that sounds right but isn't. This has a name.
Hallucination is when a model states something false as if it were fact — an invented statistic, a fake quote, a court case that never happened, a software function that doesn't exist. It isn't lying; lying needs intent. The model is doing exactly its job — writing text that fits the pattern of a correct answer — and a confident-sounding wrong answer fits that pattern just as well as a right one. The fluent tone is not evidence of anything. It's the one thing you always get, true or false.
A few more limits worth naming plainly:
It doesn't know what happened after it was trained. The model's knowledge is frozen at the point its training data was collected. Ask about yesterday's news or a price that changes hourly, and unless the app is actively looking it up and feeding it in, it's guessing from an out-of-date snapshot. (Many tools do now search the web or read your files — but that's the app adding fresh information, not the model already having it.)
It can't reliably do arithmetic or careful step-by-step reasoning on its own. It's pattern-matching what a calculation looks like, not calculating. It'll often nail an easy sum and then quietly botch a longer one. For anything that has to be exact, treat its math as a rough draft. (Here too, some tools now hand the actual math to a calculator behind the scenes — but on its own, the model is guessing.)
It has no memory of you between conversations. Each new chat starts blank. It feels like it remembers because, within one conversation, the whole visible history gets fed back in every time you hit send — that's the "memory." Close the window and, unless the app deliberately saves notes and re-feeds them, that context is gone. A big context window (how much text the model can take in at once) is not the same as the model remembering you. It's short-term working space, not a diary.
It doesn't truly understand. It has no picture of the world behind the words, no beliefs, no way to check its own output against reality. It has statistics about language. That's enough to be genuinely useful, and nowhere near enough to trust blindly.
A rule of thumb for trust
Here's a simple test for when to lean on an answer and when to check it:
If getting it wrong is cheap and you'd catch the mistake, trust it. If getting it wrong is costly and you wouldn't catch it, verify.
Rewriting your email? Cheap, and you'll read it — trust it. A brainstorm list? You're going to pick through it anyway — trust it. But a medical dose, a legal claim, a tax figure, a quote, a citation, a "fact" you're about to repeat in public? Costly, and easy to miss. Verify — against a real source, not by asking the model "are you sure?" (it'll often just confidently agree either way).
The sharpest move is to change the job. Instead of asking the model to supply facts from memory, give it the facts — paste the document, the data, the article — and ask it to work with that. You've turned a fact question, which it's bad at, into a language question, which it's great at. That one habit — matching the task to what the machine actually does — is most of what separates people who get a lot out of these tools from people who get burned by them.
One loose end: I've been saying the model predicts the next "word-piece." Those pieces have a real name — tokens — and how your words get chopped into them, before the model sees a thing, turns out to explain a surprising number of its quirks. That's next.