The beginner series · 5 min

Training vs. Inference: How an AI Goes From Expensive Science Project to Everyday Chatbot

How a model goes from expensive training run to everyday chatbot response.

Training is how a model learns: show it mountains of examples, measure how wrong it is, nudge its internal numbers, repeat — expensive, and done rarely. Inference is using the finished model on a new input, its numbers frozen. Almost every AI interaction you have — every chat, every completion — is inference, not training. That's why a model can answer you instantly yet can't actually learn your name mid-conversation unless the app keeps handing that fact back to it.

Reading promise: No assumed AI knowledge. Jargon gets translated before it gets used.

The Two Words

TRAINING
Examples/data → compute error → adjust model weights → repeat

INFERENCE
New input → frozen/trained model → prediction or generated output

NVIDIA’s current AI-inference material draws the same distinction: training adjusts model weights using data; inference applies the trained model to new inputs to generate predictions or responses.[1]

The Student Analogy

Training is the semester. Inference is the exam.

TRAINING
student studies 10,000 practice problems
gets feedback
changes what they know
repeats

INFERENCE
student receives a new problem
uses what was learned
produces an answer
During ordinary inference, the model does not normally update billions of weights just because you asked it a question.

What Happens During Pretraining?

For a language model, pretraining exposes the model to very large amounts of tokenized data and repeatedly asks it to predict masked/next tokens depending on the architecture/objective. For GPT-style autoregressive training, a simplified step looks like:

Training text:
"The capital of France is Paris."

Input:  "The capital of France is"
Target: " Paris"

Model predicts:
Paris 0.11
London 0.09
...

Target says Paris.
Loss measures how wrong the prediction was.
Backpropagation calculates how weights should change.
Optimizer updates weights.
Repeat. Millions/billions/trillions of token examples.

Loss: A Score for “How Wrong Were We?”

Training needs a numerical objective. A loss function gives the optimization process a signal. If the correct next token gets very low probability, loss is high. If the model assigns high probability to the correct token, loss is lower.

prediction
compare with target
LOSS
calculate gradients
adjust parameters in direction that tends to reduce loss

Backpropagation: The Blame Assignment System

A modern neural network may have billions of parameters. If an answer was wrong, which parameters should move, and by how much? Backpropagation uses gradients to propagate an error signal backward through the network. Google’s ML Crash Course describes backpropagation as the primary training algorithm used to adjust weights to minimize loss.[2]

You do not need calculus to keep the mental model: prediction → error → calculate which directions would reduce error → nudge weights → repeat.

Why Training Is So Expensive

Training repeatedly performs huge matrix operations, stores intermediate activations needed for gradient calculations, computes gradients, and updates parameters across enormous datasets. Large models are commonly trained in parallel across many accelerators.

TRAINING COST DRIVERS
- huge datasets
- many parameters
- repeated forward passes
- backward passes / gradients
- optimizer state
- distributed communication
- long training runs

Then the Training Run Stops

After training, you have a set of learned parameter values—a model checkpoint. You can now deploy it for inference.

DATA + COMPUTE
TRAINING
MODEL CHECKPOINT
(weights / parameters)
DEPLOYMENT
INFERENCE

What Happens When You Chat With a Model?

That is inference. Your prompt is tokenized, passed through the trained network, next-token scores are computed, a token is generated, and the process repeats. The weights are being used, not ordinarily relearned from your message.

Your prompt
tokenize
forward pass through trained model
next-token distribution
generate token
repeat until answer ends

Training vs. Fine-Tuning vs. Inference

StageWhat changes?Typical purpose
PretrainingA large set of model weightsLearn broad language/world/task patterns
Fine-tuningModel weights continue changingAdapt behavior to a narrower task/style/domain
InferenceUsually no weight updateUse the trained model on new inputs

Fine-tuning is still training. It simply starts from an already-trained model rather than random/untrained parameters.

What About RAG? Is That Training?

No. RAG changes the input context at inference time. It does not need to change the model’s weights.

TRAINING
changes model weights

RAG
changes information supplied to the model during inference

This distinction explains why you can update a company policy in a RAG index without retraining the model.

What About “Memory”? Is That Training?

Usually no. Application memory stores information externally and retrieves it into later context. The model can appear to have learned something while its weights remain unchanged.

"Remember that Project Apollo uses MySQL."
     ↓
application stores fact
     ↓ later
application retrieves fact into context
     ↓
model answers using it

No weight update required.

Inference Has Its Own Engineering Problems

Once a model is trained, deploying it efficiently becomes a different optimization problem: latency, throughput, memory footprint, batching, caching, quantization, hardware utilization, and the cost of generating each token. NVIDIA describes inference as the deployment phase where learned capability is applied to new data.[1]

TRAINING optimizes: learn good parameters
INFERENCE optimizes: serve useful outputs fast/cheap/reliably

A Cybersecurity Example

Suppose you fine-tune a model on thousands of analyst-reviewed alert classifications. That fine-tuning phase changes model parameters. Then a new alert arrives tomorrow:

NEW ALERT
WINWORD.EXE → powershell.exe -enc ...
        ↓
    INFERENCE
        ↓
model classifies "high-risk suspicious execution"

If the same system then queries Splunk for the user’s live events, that retrieval happens during inference. The model can combine learned behavior with current evidence without retraining on the new events.

The Cheat Sheet

QuestionTrainingInference
Are weights updated?YesUsually no
Uses known examples/targets?Yes, to learnTakes new inputs
Main cost shapeLarge repeated forward + backward computationForward/generation computation
User chatting with an AI?NoYes
Fine-tuning?YesNo
RAG lookup?No—it supplies contextYes, usually part of runtime
Training is how the model learns its parameters. Inference is the model spending those learned parameters on a new problem.

Sources