The beginner series · 8 min
Parameters, Weights & Neural Networks Without the Math Bullshit
What billions of parameters actually are, without requiring calculus.
A “70-billion-parameter model” means 70 billion learned numbers. A parameter — usually a weight — is just a number that controls how strongly one signal pushes on the next step of a calculation; training sets them, and stacked in layers they define the whole function the model computes. No single number stores “Paris” or “malware” — knowledge is smeared across millions of them at once. And more parameters isn't automatically smarter: data, training, and architecture decide whether the numbers are any good.
Reading promise: No assumed AI knowledge. Jargon gets translated before it gets used.
Start With One Tiny “Neuron”
Forget language models. Imagine we want a tiny model to estimate whether an email is suspicious using three features: number of links, whether the sender is new, and whether the message asks for a password reset.
inputs links = 4 new_sender = 1 password_reset = 1 Each input gets multiplied by a learned weight: 4 × w1 1 × w2 1 × w3 Then add a bias and apply a nonlinear function.
Google’s neural-network material describes network parameters as learned weights and biases; even a tiny hidden layer can contain multiple weight and bias parameters.[1]
A Weight Is Just a Number
Suppose our learned weights were:
w_links = 0.2 w_new_sender = 1.1 w_password_reset = 1.8 bias = -0.7
Those numbers determine how strongly each input influences the next computation. During training, the optimization process changes them.
Parameter is the broad term. Weights and biases are common kinds of parameters.
A Network Is Lots of These Operations Connected Together
INPUTS
x1 x2 x3
\ | /
\ | /
HIDDEN LAYER
o o o o
\ | / \ |
ANOTHER LAYER
o o
\ /
OUTPUTReal neural networks are usually represented as matrix operations rather than literally iterating cartoon neurons one by one. But the diagram captures the idea: layers transform numbers into new numbers using learned parameters.
Why “Neural” Network?
The terminology was historically inspired by biological neurons, but modern artificial neural networks are mathematical systems, not simulations of a human brain neuron-for-neuron. Treat the biology analogy lightly.
What Does a Parameter Count Mean?
When someone says a model has 7 billion, 70 billion, or hundreds of billions of parameters, they mean the network contains that many learned numeric values (or values counted under the architecture’s parameter accounting). Those numbers define the function the model computes.
TINY MODEL 21 parameters LARGE LANGUAGE MODEL 7,000,000,000+ parameters Same basic category: learned numeric values. Wildly different scale and architecture.
Google’s worked neural-network example explicitly counts weights plus biases to derive a model’s parameter total.[1]
Where Is the Knowledge?
This question invites a bad mental model:
Parameter 1 = Paris Parameter 2 = France Parameter 3 = dog Parameter 4 = malware
That is not how it works. Useful concepts are represented distributedly through many interacting parameters and activations. A specific factual or behavioral capability may depend on patterns spanning many layers and weights.
The model is closer to an enormous learned function than an enormous spreadsheet of facts.
Weights vs. Activations
Another distinction: weights are relatively persistent learned parameters. Activations are temporary values produced while the network processes a particular input.
WEIGHTS learned during training reused across requests ACTIVATIONS computed for this specific prompt change from request to request
If you ask about a phishing email and then ask about Shakespeare, the model’s weights are largely the same; the internal activations produced by the two inputs are different.
Why Do We Need Activation Functions?
If every layer only performed simple linear transformations, stacking many layers would collapse mathematically into another linear transformation. Nonlinear activation functions let neural networks model far more complicated relationships. Google’s ML material lists nodes, learned weights/biases, and activation functions as core pieces of neural-network layers.[2]
Training Is Weight Adjustment
1. model sees an example 2. model predicts 3. calculate loss 4. backpropagation computes gradients 5. optimizer nudges weights/biases 6. repeat
Google describes gradient descent as an iterative technique for finding weight and bias values that reduce loss.[3] In a large neural network the geometry is vastly more complex, but the intuition survives: use error signals to improve parameters.
Parameters vs. Hyperparameters
These are frequently confused.
| Term | Who/what sets it? | Example |
|---|---|---|
| Parameter | Learned by the training process | weight = 0.127, bias = -0.04 |
| Hyperparameter | Chosen/configured around training or architecture | learning rate, batch size, number of epochs |
Google’s ML Crash Course identifies learning rate, batch size, and epochs as hyperparameters that influence training rather than values learned as ordinary model weights.[4]
Why More Parameters Can Help
More parameters give a model more representational capacity, but parameter count by itself is not a quality score. Data quality, training objective, architecture, optimization, context length, post-training, tool access, and inference-time reasoning all matter. A poorly trained larger model can lose to a better-designed smaller one on a given task.
Do not buy the “more billions = automatically smarter” marketing shortcut.
How Parameters Take Up Memory
A parameter is stored numerically. The number format matters. If you stored 70 billion parameters at 16 bits each, the raw parameter storage alone is roughly 140 GB before additional runtime overhead. Lower-precision representations and quantization can reduce memory use.
70 billion parameters × 2 bytes ≈ 140 GB This is only a back-of-the-envelope parameter-weight calculation. Inference also needs other memory: caches, activations, runtime buffers, etc.
Why GPUs Are So Useful
Neural networks spend enormous amounts of time doing matrix and tensor arithmetic. GPUs and other accelerators can perform many mathematical operations in parallel, which is why they became central to large-scale neural-network training and inference.
Where Does Attention Fit?
The Transformer from Article 6 is still a neural network. Attention itself contains learned projection weights. Feed-forward sublayers contain more learned weights. Embedding tables contain learned values. “Transformer” describes an architecture built from neural-network operations—it does not replace the concept of parameters.
LLM PARAMETERS INCLUDE THINGS LIKE - token embedding matrices - attention projection matrices - feed-forward layer matrices - normalization parameters - output projection / related weights Exact architecture varies by model.
A Cybersecurity Analogy That Actually Works
Imagine a detection rule:
IF parent = WINWORD.EXE +3 risk IF child = powershell.exe +4 risk IF encoded command +3 risk IF known admin script -5 risk score ≥ 6 → investigate
Those hand-written risk values are a crude analogy to weights: they determine influence. A neural network learns vastly larger collections of such numeric relationships automatically rather than a human writing +3 next to every condition. Unlike the toy rule, the learned relationships are distributed, layered, nonlinear, and difficult to interpret directly.
The Entire Story So Far
TRAINING DATA
↓
TOKENIZATION
↓
NEURAL NETWORK with many PARAMETERS
↓
PREDICTION
↓
LOSS
↓
BACKPROPAGATION + OPTIMIZER
↓
adjust WEIGHTS
↺ repeat
After training:
PROMPT → trained network → next-token predictions → responseThe Cheat Sheet
| Term | Plain English |
|---|---|
| Parameter | A learned numeric value in the model |
| Weight | A common parameter controlling how strongly information contributes to a computation |
| Bias | A learned offset term |
| Activation | A temporary value produced while processing a particular input |
| Layer | A stage of neural-network transformation |
| Activation function | A nonlinear transformation that increases what the network can represent |
| Gradient | Information about how changing parameters would change loss |
| Backpropagation | Method for computing those gradients through the network |
| Hyperparameter | A training/configuration choice rather than an ordinary learned weight |
A billion-parameter model is, at a very high level, a machine whose behavior is governed by roughly a billion learned numbers interacting through a particular architecture.
Sources
- [1] Google Machine Learning Crash Course — Neural networks: Nodes and hidden layers. — https://developers.google.com/machine-learning/crash-course/neural-networks/nodes-hidden-layers
- [2] Google Machine Learning Crash Course — Activation functions. — https://developers.google.com/machine-learning/crash-course/neural-networks/activation-functions
- [3] Google Machine Learning Crash Course — Gradient descent. — https://developers.google.com/machine-learning/crash-course/linear-regression/gradient-descent
- [4] Google Machine Learning Crash Course — Hyperparameters. — https://developers.google.com/machine-learning/crash-course/linear-regression/hyperparameters
- [5] Vaswani et al. — Attention Is All You Need. — https://arxiv.org/abs/1706.03762
Series Map: How the Pieces Fit Together
USER
│
▼
AGENT
│
┌─────┴─────┐
▼ ▼
PROMPT TOOLS
│ │
│ MCP
│ │
│ external systems
│
┌────────┴─────────┐
▼ ▼
MEMORY / RAG live tool results
│ │
└────────┬─────────┘
▼
CONTEXT WINDOW
│
▼
TOKENS
│
▼
LLM
(neural network)
│
parameters / weights
│
▼
next-token output
TRAINING changes parameters.
INFERENCE uses them.
FINE-TUNING continues training for narrower behavior.Where the Beginner Series Can Go Next
| Proposed article | Question it answers |
|---|---|
| 9. WTF Is Attention? | How can every token “look at” other tokens, and what are Q, K, and V really doing? |
| 10. MoE, Quantization & Distillation | How do people make giant models cheaper and faster? |
| 11. Reasoning Models vs. “Normal” LLMs | What changes when a model spends more inference compute on a problem? |
| 12. Evals: How Do You Know an AI System Is Actually Good? | How to measure quality instead of judging five cherry-picked demos. |
| 13. Hallucinations: Why They Happen | Why confident falsehood is a natural failure mode of generative models. |
| 14. Tool Calling vs. Computer Use | API/tool access versus literally operating a graphical interface. |
| 15. AI Security 101 | Prompt injection, data leakage, excessive agency, supply-chain risk, and permission boundaries. |
Editorial Note
This collection intentionally favors durable mental models over vendor-specific implementation details. Where protocol or product details are version-sensitive—especially MCP—the source list points to the dated/current specification so the implementation facts can be rechecked before publication.