The beginner series · 8 min

Parameters, Weights & Neural Networks Without the Math Bullshit

What billions of parameters actually are, without requiring calculus.

A “70-billion-parameter model” means 70 billion learned numbers. A parameter — usually a weight — is just a number that controls how strongly one signal pushes on the next step of a calculation; training sets them, and stacked in layers they define the whole function the model computes. No single number stores “Paris” or “malware” — knowledge is smeared across millions of them at once. And more parameters isn't automatically smarter: data, training, and architecture decide whether the numbers are any good.

Reading promise: No assumed AI knowledge. Jargon gets translated before it gets used.

Start With One Tiny “Neuron”

Forget language models. Imagine we want a tiny model to estimate whether an email is suspicious using three features: number of links, whether the sender is new, and whether the message asks for a password reset.

inputs
links = 4
new_sender = 1
password_reset = 1

Each input gets multiplied by a learned weight:
4 × w1
1 × w2
1 × w3

Then add a bias and apply a nonlinear function.

Google’s neural-network material describes network parameters as learned weights and biases; even a tiny hidden layer can contain multiple weight and bias parameters.[1]

A Weight Is Just a Number

Suppose our learned weights were:

w_links          = 0.2
w_new_sender     = 1.1
w_password_reset = 1.8
bias             = -0.7

Those numbers determine how strongly each input influences the next computation. During training, the optimization process changes them.

Parameter is the broad term. Weights and biases are common kinds of parameters.

A Network Is Lots of These Operations Connected Together

INPUTS
 x1  x2  x3
  \  |  /
   \ | /
 HIDDEN LAYER
  o  o  o  o
   \ | / \ |
  ANOTHER LAYER
    o   o
     \ /
   OUTPUT

Real neural networks are usually represented as matrix operations rather than literally iterating cartoon neurons one by one. But the diagram captures the idea: layers transform numbers into new numbers using learned parameters.

Why “Neural” Network?

The terminology was historically inspired by biological neurons, but modern artificial neural networks are mathematical systems, not simulations of a human brain neuron-for-neuron. Treat the biology analogy lightly.

What Does a Parameter Count Mean?

When someone says a model has 7 billion, 70 billion, or hundreds of billions of parameters, they mean the network contains that many learned numeric values (or values counted under the architecture’s parameter accounting). Those numbers define the function the model computes.

TINY MODEL
21 parameters

LARGE LANGUAGE MODEL
7,000,000,000+ parameters

Same basic category: learned numeric values.
Wildly different scale and architecture.

Google’s worked neural-network example explicitly counts weights plus biases to derive a model’s parameter total.[1]

Where Is the Knowledge?

This question invites a bad mental model:

Parameter 1 = Paris
Parameter 2 = France
Parameter 3 = dog
Parameter 4 = malware

That is not how it works. Useful concepts are represented distributedly through many interacting parameters and activations. A specific factual or behavioral capability may depend on patterns spanning many layers and weights.

The model is closer to an enormous learned function than an enormous spreadsheet of facts.

Weights vs. Activations

Another distinction: weights are relatively persistent learned parameters. Activations are temporary values produced while the network processes a particular input.

WEIGHTS
learned during training
reused across requests

ACTIVATIONS
computed for this specific prompt
change from request to request

If you ask about a phishing email and then ask about Shakespeare, the model’s weights are largely the same; the internal activations produced by the two inputs are different.

Why Do We Need Activation Functions?

If every layer only performed simple linear transformations, stacking many layers would collapse mathematically into another linear transformation. Nonlinear activation functions let neural networks model far more complicated relationships. Google’s ML material lists nodes, learned weights/biases, and activation functions as core pieces of neural-network layers.[2]

Training Is Weight Adjustment

1. model sees an example
2. model predicts
3. calculate loss
4. backpropagation computes gradients
5. optimizer nudges weights/biases
6. repeat

Google describes gradient descent as an iterative technique for finding weight and bias values that reduce loss.[3] In a large neural network the geometry is vastly more complex, but the intuition survives: use error signals to improve parameters.

Parameters vs. Hyperparameters

These are frequently confused.

TermWho/what sets it?Example
ParameterLearned by the training processweight = 0.127, bias = -0.04
HyperparameterChosen/configured around training or architecturelearning rate, batch size, number of epochs

Google’s ML Crash Course identifies learning rate, batch size, and epochs as hyperparameters that influence training rather than values learned as ordinary model weights.[4]

Why More Parameters Can Help

More parameters give a model more representational capacity, but parameter count by itself is not a quality score. Data quality, training objective, architecture, optimization, context length, post-training, tool access, and inference-time reasoning all matter. A poorly trained larger model can lose to a better-designed smaller one on a given task.

Do not buy the “more billions = automatically smarter” marketing shortcut.

How Parameters Take Up Memory

A parameter is stored numerically. The number format matters. If you stored 70 billion parameters at 16 bits each, the raw parameter storage alone is roughly 140 GB before additional runtime overhead. Lower-precision representations and quantization can reduce memory use.

70 billion parameters × 2 bytes ≈ 140 GB

This is only a back-of-the-envelope parameter-weight calculation.
Inference also needs other memory: caches, activations, runtime buffers, etc.

Why GPUs Are So Useful

Neural networks spend enormous amounts of time doing matrix and tensor arithmetic. GPUs and other accelerators can perform many mathematical operations in parallel, which is why they became central to large-scale neural-network training and inference.

Where Does Attention Fit?

The Transformer from Article 6 is still a neural network. Attention itself contains learned projection weights. Feed-forward sublayers contain more learned weights. Embedding tables contain learned values. “Transformer” describes an architecture built from neural-network operations—it does not replace the concept of parameters.

LLM PARAMETERS INCLUDE THINGS LIKE
- token embedding matrices
- attention projection matrices
- feed-forward layer matrices
- normalization parameters
- output projection / related weights

Exact architecture varies by model.

A Cybersecurity Analogy That Actually Works

Imagine a detection rule:

IF parent = WINWORD.EXE       +3 risk
IF child = powershell.exe      +4 risk
IF encoded command             +3 risk
IF known admin script          -5 risk

score ≥ 6 → investigate

Those hand-written risk values are a crude analogy to weights: they determine influence. A neural network learns vastly larger collections of such numeric relationships automatically rather than a human writing +3 next to every condition. Unlike the toy rule, the learned relationships are distributed, layered, nonlinear, and difficult to interpret directly.

The Entire Story So Far

TRAINING DATA
     ↓
TOKENIZATION
     ↓
NEURAL NETWORK with many PARAMETERS
     ↓
PREDICTION
     ↓
LOSS
     ↓
BACKPROPAGATION + OPTIMIZER
     ↓
adjust WEIGHTS
     ↺ repeat

After training:
PROMPT → trained network → next-token predictions → response

The Cheat Sheet

TermPlain English
ParameterA learned numeric value in the model
WeightA common parameter controlling how strongly information contributes to a computation
BiasA learned offset term
ActivationA temporary value produced while processing a particular input
LayerA stage of neural-network transformation
Activation functionA nonlinear transformation that increases what the network can represent
GradientInformation about how changing parameters would change loss
BackpropagationMethod for computing those gradients through the network
HyperparameterA training/configuration choice rather than an ordinary learned weight
A billion-parameter model is, at a very high level, a machine whose behavior is governed by roughly a billion learned numbers interacting through a particular architecture.

Sources

Series Map: How the Pieces Fit Together

                           USER
                            │
                            ▼
                         AGENT
                            │
                      ┌─────┴─────┐
                      ▼           ▼
                    PROMPT       TOOLS
                      │           │
                      │          MCP
                      │           │
                      │      external systems
                      │
             ┌────────┴─────────┐
             ▼                  ▼
          MEMORY / RAG       live tool results
             │                  │
             └────────┬─────────┘
                      ▼
                 CONTEXT WINDOW
                      │
                      ▼
                    TOKENS
                      │
                      ▼
                     LLM
              (neural network)
                      │
           parameters / weights
                      │
                      ▼
              next-token output

TRAINING changes parameters.
INFERENCE uses them.
FINE-TUNING continues training for narrower behavior.

Where the Beginner Series Can Go Next

Proposed articleQuestion it answers
9. WTF Is Attention?How can every token “look at” other tokens, and what are Q, K, and V really doing?
10. MoE, Quantization & DistillationHow do people make giant models cheaper and faster?
11. Reasoning Models vs. “Normal” LLMsWhat changes when a model spends more inference compute on a problem?
12. Evals: How Do You Know an AI System Is Actually Good?How to measure quality instead of judging five cherry-picked demos.
13. Hallucinations: Why They HappenWhy confident falsehood is a natural failure mode of generative models.
14. Tool Calling vs. Computer UseAPI/tool access versus literally operating a graphical interface.
15. AI Security 101Prompt injection, data leakage, excessive agency, supply-chain risk, and permission boundaries.

Editorial Note

This collection intentionally favors durable mental models over vendor-specific implementation details. Where protocol or product details are version-sensitive—especially MCP—the source list points to the dated/current specification so the implementation facts can be rechecked before publication.

Parameters, Weights & Neural Networks Without the Math Bullshit — WTF Is…? — Plain-English AI · AdversariaLLM