The beginner series · 5 min
Training vs. Inference: How an AI Goes From Expensive Science Project to Everyday Chatbot
How a model goes from expensive training run to everyday chatbot response.
Training is how a model learns: show it mountains of examples, measure how wrong it is, nudge its internal numbers, repeat — expensive, and done rarely. Inference is using the finished model on a new input, its numbers frozen. Almost every AI interaction you have — every chat, every completion — is inference, not training. That's why a model can answer you instantly yet can't actually learn your name mid-conversation unless the app keeps handing that fact back to it.
Reading promise: No assumed AI knowledge. Jargon gets translated before it gets used.
The Two Words
TRAINING Examples/data → compute error → adjust model weights → repeat INFERENCE New input → frozen/trained model → prediction or generated output
NVIDIA’s current AI-inference material draws the same distinction: training adjusts model weights using data; inference applies the trained model to new inputs to generate predictions or responses.[1]
The Student Analogy
Training is the semester. Inference is the exam.
TRAINING student studies 10,000 practice problems gets feedback changes what they know repeats INFERENCE student receives a new problem uses what was learned produces an answer
During ordinary inference, the model does not normally update billions of weights just because you asked it a question.
What Happens During Pretraining?
For a language model, pretraining exposes the model to very large amounts of tokenized data and repeatedly asks it to predict masked/next tokens depending on the architecture/objective. For GPT-style autoregressive training, a simplified step looks like:
Training text: "The capital of France is Paris." Input: "The capital of France is" Target: " Paris" Model predicts: Paris 0.11 London 0.09 ... Target says Paris. Loss measures how wrong the prediction was. Backpropagation calculates how weights should change. Optimizer updates weights. Repeat. Millions/billions/trillions of token examples.
Loss: A Score for “How Wrong Were We?”
Training needs a numerical objective. A loss function gives the optimization process a signal. If the correct next token gets very low probability, loss is high. If the model assigns high probability to the correct token, loss is lower.
Backpropagation: The Blame Assignment System
A modern neural network may have billions of parameters. If an answer was wrong, which parameters should move, and by how much? Backpropagation uses gradients to propagate an error signal backward through the network. Google’s ML Crash Course describes backpropagation as the primary training algorithm used to adjust weights to minimize loss.[2]
You do not need calculus to keep the mental model: prediction → error → calculate which directions would reduce error → nudge weights → repeat.
Why Training Is So Expensive
Training repeatedly performs huge matrix operations, stores intermediate activations needed for gradient calculations, computes gradients, and updates parameters across enormous datasets. Large models are commonly trained in parallel across many accelerators.
TRAINING COST DRIVERS - huge datasets - many parameters - repeated forward passes - backward passes / gradients - optimizer state - distributed communication - long training runs
Then the Training Run Stops
After training, you have a set of learned parameter values—a model checkpoint. You can now deploy it for inference.
What Happens When You Chat With a Model?
That is inference. Your prompt is tokenized, passed through the trained network, next-token scores are computed, a token is generated, and the process repeats. The weights are being used, not ordinarily relearned from your message.
Training vs. Fine-Tuning vs. Inference
| Stage | What changes? | Typical purpose |
|---|---|---|
| Pretraining | A large set of model weights | Learn broad language/world/task patterns |
| Fine-tuning | Model weights continue changing | Adapt behavior to a narrower task/style/domain |
| Inference | Usually no weight update | Use the trained model on new inputs |
Fine-tuning is still training. It simply starts from an already-trained model rather than random/untrained parameters.
What About RAG? Is That Training?
No. RAG changes the input context at inference time. It does not need to change the model’s weights.
TRAINING changes model weights RAG changes information supplied to the model during inference
This distinction explains why you can update a company policy in a RAG index without retraining the model.
What About “Memory”? Is That Training?
Usually no. Application memory stores information externally and retrieves it into later context. The model can appear to have learned something while its weights remain unchanged.
"Remember that Project Apollo uses MySQL."
↓
application stores fact
↓ later
application retrieves fact into context
↓
model answers using it
No weight update required.Inference Has Its Own Engineering Problems
Once a model is trained, deploying it efficiently becomes a different optimization problem: latency, throughput, memory footprint, batching, caching, quantization, hardware utilization, and the cost of generating each token. NVIDIA describes inference as the deployment phase where learned capability is applied to new data.[1]
TRAINING optimizes: learn good parameters INFERENCE optimizes: serve useful outputs fast/cheap/reliably
A Cybersecurity Example
Suppose you fine-tune a model on thousands of analyst-reviewed alert classifications. That fine-tuning phase changes model parameters. Then a new alert arrives tomorrow:
NEW ALERT
WINWORD.EXE → powershell.exe -enc ...
↓
INFERENCE
↓
model classifies "high-risk suspicious execution"If the same system then queries Splunk for the user’s live events, that retrieval happens during inference. The model can combine learned behavior with current evidence without retraining on the new events.
The Cheat Sheet
| Question | Training | Inference |
|---|---|---|
| Are weights updated? | Yes | Usually no |
| Uses known examples/targets? | Yes, to learn | Takes new inputs |
| Main cost shape | Large repeated forward + backward computation | Forward/generation computation |
| User chatting with an AI? | No | Yes |
| Fine-tuning? | Yes | No |
| RAG lookup? | No—it supplies context | Yes, usually part of runtime |
Training is how the model learns its parameters. Inference is the model spending those learned parameters on a new problem.
Sources
- [1] NVIDIA — What Is AI Inference?. — https://www.nvidia.com/en-us/glossary/ai-inference/
- [2] Google Machine Learning Crash Course — Neural networks: Backpropagation. — https://developers.google.com/machine-learning/crash-course/neural-networks/backpropagation
- [3] Google Machine Learning Crash Course — Linear regression: Gradient descent. — https://developers.google.com/machine-learning/crash-course/linear-regression/gradient-descent
- [4] NVIDIA — What Is AI Training?. — https://www.nvidia.com/en-us/glossary/ai-training/