All About LLMs, from AI to Z
A practical guide on how large language models actually work — starting from plain-language basics, then tokens and attention, RAG, agents, AI security, evaluation, and serving.
11 modules · 80 lessons
On this page
Start Here — LLMs in Plain Language
Brand new to this? Start here. Five short, plain-language lessons — what a large language model actually is, what it can and can't do, how it runs the moment you hit send, and how it was built — so the rest of the course opens the hood on something you already understand.
A one-page plain-language glossary you can keep: every core term — model, parameters, prompt, token, context window, inference, training, hallucination — in a single sentence each.
- What a large language model actually is5 minA large language model is autocomplete trained on a huge slice of the written world — it predicts the next chunk of text, over and over, and that one trick is where all its power and all its limits come from.
- What LLMs are good at — and what they are not5 minLLMs are fluent-language machines: great at reshaping words, unreliable about facts — and knowing the difference is the whole skill.
- How a model runs: from your message to its reply5 minFollow a single message from the moment you hit send to the reply that streams back, and you'll understand what a language model is actually doing the whole time.
- How a model learns: pretraining, then made helpful5 minA model is built in two stages — first it reads a mountain of text to predict what comes next, then it's taught to be helpful — and knowing that explains why it can be brilliant and wrong in the same breath.
- The words everyone uses: a plain-language glossary7 minThe handful of words you'll meet everywhere in AI — model, weights, prompt, token, context, inference, training, hallucination, and "reasoning" — explained through one short chat so they connect instead of float.
Foundational Mechanics
How a language model actually turns text into tokens, predicts the next one, and why the context window is a budget, not a memory.
A notebook that tokenizes text, visualizes attention dimensions, measures context growth, and explains the memory footprint of a single inference request.
- How tokenization works (and why it bites you)7 minYour text is compiled to integers by a frozen, learned lookup — and every surprise in cost, context, and safety starts at that boundary.
- Embeddings: turning tokens into geometry7 minHow language models turn discrete tokens into dense vectors where distance means similarity — and where that geometry quietly lies to you.
- The Transformer block, one mental model6 minA transformer is one running per-token vector that every block reads from and adds back to; hold that picture and the rest follows.
- Self-attention, actually explained6 minSelf-attention is a soft, differentiable dictionary lookup where every key matches a little — the source of both its power and its quadratic cost.
- Why order needs to be encoded7 minSelf-attention is blind to word order by construction; four positional schemes put order back in, and each one breaks differently past its training length.
- Decoding: how the next token is chosen7 minThe decoder collapses a vocabulary-sized score vector into one token: softmax, temperature, top-k/top-p, and why temperature 0 still isn't deterministic.
- KV caching: why generation is one token at a time7 minAutoregressive decode forces one token per forward pass; the KV cache turns the quadratic wall linear but relocates the ceiling from compute onto memory bandwidth.
- The Context Window Is Not Memory7 minA model re-reads the whole transcript every turn from a finite, costly, unevenly-used window — and mistaking that for memory breeds production bugs.
Model Families & Architectures
Encoder, decoder, and encoder-decoder designs, the open-weight landscape, and how to choose a model family on purpose instead of by leaderboard.
A written comparison that selects a candidate model family for a concrete task and defends the choice on architecture, licence, and fit.
- Three Architectures: Encoder-Only, Decoder-Only, Encoder-Decoder6 minThe transformer split into three families along two axes, attention masking and training objective, and those two choices explain what each model can and can't do.
- Why Decoder-Only Won for Generation6 minFew-shot prompting turned every NLP task into next-token prediction, retiring per-task fine-tuning and making one decoder cheaper than a zoo of encoders.
- Mixture-of-Experts: Total vs Active Parameters6 minSparse routing sends each token to a few experts of many, making an MoE compute-cheap per token yet still full-price in memory to hold every expert.
- Small Language Models: When Smaller Wins6 minWhere sub-frontier models beat the biggest one on latency, cost, edge, and narrow accuracy, and how data quality and distillation move the scaling curve left.
- Multimodality: Bolting Vision and Audio onto a Text Model6 minA text-only LLM gains eyes and ears when a frozen encoder turns pixels or audio into vectors and a small trained projector maps them into the model's token space.
- Model Routing and the Real Selection Tradeoffs6 minHow to route between models at runtime across capability, latency, cost, and context — with cascades, difficulty routing, and the ways routers break.
- Open-Weight vs Closed Families: What It Means in Practice6 minOpen-weight and closed API models differ in data control, licensing, fine-tuning access, and cost shape, and that sourcing choice is separate from runtime routing.
Data, Pre-Training & Post-Training
Where a model's behavior comes from: the pre-training corpus, then the fine-tuning, RLHF, and quantization that shape and shrink it.
Fine-tune a small model with LoRA/QLoRA, evaluate the base against the adapted version, quantize for inference, and produce a model card.
- Building the Corpus: Crawl, Dedup, Filter — and Poison6 minHow a pre-training corpus is really built from Common Crawl: extraction, MinHash dedup, quality classifiers, and why poisoning 0.01% of it costs about $60 a year.
- What Pre-Training Actually Optimizes: Next-Token Prediction7 minPre-training minimizes one loss, cross-entropy on the next token, and that single objective explains perplexity, base models, and why raw prediction yields broad capability.
- Scaling Laws: Why Chinchilla Says Data Is the Binding Constraint6 minHow compute-optimal scaling works, why Kaplan 2020 pushed the field toward oversized undertrained models, and why Chinchilla reframed data as the real bottleneck.
- From Base Model to Assistant: Supervised Fine-Tuning6 minHow a small set of hand-written demonstration pairs and the ordinary next-token loss turn a text-completing base model into one that follows instructions.
- Aligning to Preferences: RLHF and the Simpler DPO7 minHow preference optimization turns human A/B judgments into model behavior: the reward-model-plus-PPO pipeline, and DPO's collapse of it into one classification loss.
- Fine-Tuning on a Budget: LoRA and QLoRA6 minWhy full fine-tuning burns memory you don't need to spend, and how low-rank adapters plus 4-bit quantization put a 65B tune on one GPU.
- Shrinking Weights for Inference: GPTQ, AWQ, and GGUF7 minPost-training quantization drops weights to 4-bit for a few percent of quality, and GGUF is a shipping container, not a compression method.
- Teaching Small Models from Big Ones: Distillation6 minHow distillation trains a small student to copy a big teacher's output distribution — the original soft-label method, and what today's "distilled" open models actually do.
RAG & Knowledge Systems
Grounding a model in your own documents: chunking, embeddings, hybrid retrieval, reranking, provenance, and access-control-aware knowledge.
An enterprise-style RAG application with hybrid retrieval, reranking, provenance, access-control metadata, citations, and an evaluation dataset.
- Why RAG Exists: Parametric Weights vs. Retrieved Context6 minTwo hard limits — a fixed context window and weights frozen at training cutoff — force a split between knowledge baked into weights and a corpus fetched at inference.
- Ingestion and Chunking: Turning Documents into Retrievable Units6 minHow raw PDFs, HTML, code, and tickets become indexable chunks, and why the boundaries you draw at ingest quietly cap everything retrieval can ever do.
- Embeddings as a Search Index: Dense Vectors and Approximate Nearest Neighbors7 minHow an embedding model becomes a retrieval index, why brute-force cosine search stops scaling, and how HNSW and FAISS trade a sliver of recall for orders-of-magnitude speed.
- Hybrid Retrieval: Why Lexical BM25 Still Beats Embeddings Sometimes6 minBM25 ranks by exact-token statistics, so it nails rare identifiers and error codes that dense vectors blur; here's how to fuse it with embeddings and cover both failure modes.
- Reranking with Cross-Encoders: Precision on the Shortlist6 minFirst-stage retrieval encodes query and document apart for speed; a cross-encoder reads them together to fix the ranking on the top-k.
- Context Assembly, Citation, and Lost in the Middle6 minHow retrieved chunks become the real prompt: budgeting, dedup, provenance, and why rank position is a correctness lever rather than cosmetics.
- GraphRAG: Retrieval Over Entities and Relationships6 minFlat chunk retrieval breaks on multi-hop and whole-corpus questions; GraphRAG indexes an entity graph and community summaries so query time becomes a walk, not a lucky guess.
- Access Control: The Retrieval Layer That Leaks Data6 minA retriever that searches the whole corpus will hand a user documents they were never allowed to see. Here is why, and how to enforce permissions at query time.
Agents & Tool Integration
What an agent really is — a loop, not magic — plus function calling, MCP, memory, planning, multi-agent orchestration, and the confused deputy.
A tool-using workflow that can read data freely but requires deterministic permission checks and explicit approval before any consequential write.
- What an Agent Actually Is: Loop, Not Magic6 minAn agent is five ordinary parts wired into a loop; "autonomous" describes how much authority that wiring hands the model, not a capability hidden in the weights.
- Tool Interfaces and Function Calling6 minHow a model emits a schema-constrained tool-call request, why the runtime (not the model) executes it, and where that boundary leaks.
- MCP: A Standard Tool and Resource Interface6 minMCP turns per-vendor function-calling into a shared client/server protocol. Here's how its transport, primitives, and trust boundaries actually work.
- State and Memory: Context Window vs. Long-Term Store7 minAn agent has no memory of its own. It has a context window the loop rebuilds each step, and a store you engineer to refill it.
- Planning and the Control Loop: ReAct, Plan-then-Execute, Reflection6 minThree ways an agent picks its next action and recovers from mistakes — interleaved reason/act, upfront planning, and self-critique — read as policy choices over one loop.
- Multi-Agent Orchestration: When It Helps, When It's Overkill6 minHow to tell real task decomposition from cargo-culted agent sprawl, using honest numbers on cost, latency, and coordination failure.
- The Load-Bearing Truth: Danger Is Permission, Not Intelligence7 minAn agent's blast radius is set by the authority you hand its tools, not by how smart the model is — confused-deputy theory explains why.
AI Cybersecurity
The LLM threat model in practice: prompt injection direct and indirect, jailbreaks, insecure output handling, data exfiltration, and poisoning.
Threat-model and attack a synthetic LLM application to discover its realistic failure paths before an adversary does.
- The LLM Application Threat Model6 minInstructions and untrusted data share one token channel, so the classic code/data boundary collapses. Here's the resulting attack surface, mapped to OWASP LLM Top 10 and NIST AI RMF.
- Prompt Injection: Direct and Indirect6 minWhy an LLM can't tell your instructions from an attacker's text, how the payload arrives from documents nobody read, and why fencing the prompt fixes nothing.
- Jailbreaks vs. Prompt Injection7 minPeople constantly conflate these two attacks: one defeats a model's safety training, the other hijacks your app's trusted instructions. Different owners, different fixes.
- Insecure Output Handling6 minModel output is untrusted input to whatever runs it next: render sinks become XSS, generated URLs become SSRF, generated code becomes injection.
- Data Exfiltration Through the Model6 minInjected instructions turn an LLM into an exfiltration channel, leaking system prompts, RAG data, and other users' data through markdown image beacons and attacker-directed tool calls.
- Agent Abuse and the Confused Deputy6 minTool-using agents get talked into misusing privileges they legitimately hold. That is the confused-deputy pattern, and excessive agency sets the blast radius.
- Training-Time Poisoning and the Model Supply Chain7 minThe attacks that land before a single token is generated: backdoored weights, poisoned corpora, and malicious model-hub artifacts you never built.
- Identity and Authentication for AI Systems6 minWhy an LLM agent is the perfect confused deputy, and how scoping credentials to the requesting user keeps one injection from becoming one breach.
Defense & Governance
Reducing blast radius: trust boundaries, least privilege, an authorization layer around tools, guardrails and their limits, sandboxing, and red-teaming your own AI.
Apply least privilege, policy gates, human approvals, and guardrails to the synthetic application and prove each control with a measurement.
- Trust Boundaries: Why the Prompt Is Not a Security Control6 minThe context window is untrusted input with no enforced channel separation, so security belongs in deterministic code around the model, not in system-prompt instructions.
- Least Privilege for Agents: Scoping Tools and Credentials7 minAn over-privileged agent acting on attacker-controlled input is a confused deputy; scoping tools, credentials, and identity to the end user shuts it down.
- Doing Security in Code: An Authorization Layer Around Tool Calls6 minPut a deterministic policy check between the model's tool call and its execution, so permissions never depend on the model choosing to obey.
- Guardrails: Input/Output Filtering and Why a Filter Is Not a Boundary6 minInput and output guardrails like Llama Guard and NeMo cut LLM risk, but they are probabilistic classifiers layered on top of a code-enforced authorization boundary, never a replacement for it.
- Sandboxing: Jailing Tool Execution and Model-Generated Code7 minHow to contain the blast radius when an agent runs code or calls tools: pick an isolation boundary, lock down filesystem and egress, drop privileges, treat every model output as hostile.
- Human-in-the-Loop: Gating Irreversible and High-Stakes Actions6 minWhen an agent can move money, delete data, or send mail, the approval gate is a real security control, but only if it fights rubber-stamping instead of manufacturing it.
- Red-Teaming: Attacking Your Own AI Before Someone Else Does7 minA working method for attacking your own LLM system: jailbreak and injection suites, tool-abuse probes, indirect injection via retrieval, with garak, PyRIT, and a regression harness.
- Governance and Compliance: NIST AI RMF, the EU AI Act, and Provenance6 minA working model of AI governance: the NIST RMF as an operating loop, the EU AI Act's risk tiers and obligations, and the provenance artifacts that make a system auditable.
Evaluation, Benchmarking & Observability
How to measure an LLM system instead of guessing — from model benchmarks up through RAG and agent evaluation, judges, and production traces.
A CI-compatible evaluation harness that fails a build when quality, security, latency, or cost thresholds regress.
- The Evaluation Hierarchy: Five Layers of "Is It Good?"7 minA benchmark score tells you almost nothing about whether your app works. This lays out the ladder that does — from deterministic component tests up through model, RAG, and agent eval to the business metrics that are the only real verdict.
- Reading a Benchmark Without Getting Fooled7 minMMLU and SWE-bench measure two different philosophies of "smart." What each actually tests, why training contamination quietly inflates scores, and five questions that tell you whether a number transfers to your problem.
- Building a Production Evaluation Suite7 minGolden-set, adversarial, and regression cases, plus slice analysis — the four things a real eval suite has that a benchmark does not, and why the suite grows with every incident.
- LLM-as-a-Judge as a Measurement Instrument7 minUsing a model to grade a model is fine — as long as you treat the judge like any instrument: calibrate it against human labels, probe its biases, version it into every number it produces, and re-check it when it drifts.
- Evaluating RAG: Measure Retrieval and Generation Apart7 minIf you only score the final answer, you cannot tell a retrieval miss from a generation lie. Decompose the pipeline and each failure gets its own number.
- Agent Evaluation: Beyond the Final Answer7 minA correct answer reached through an unauthorized action should FAIL. How to score the trajectory an agent took — tool choice, unauthorized reaches, recovery, cost, side effects — not just where it landed.
- Observability: Reading the Trace7 minAn agent run is an execution graph — model calls, retrievals, and tool calls, each with its own numbers. Instrument it with a shared convention and debugging stops being archaeology.
- The CI Evaluation Harness: Fail the Build on Regression7 minThe module's capstone: wire your golden, adversarial, security, latency, and cost checks into CI so a quality or safety regression turns the build red and blocks the merge before it reaches users.
Infrastructure, Hardware & Production Deployment
The performance mechanics of serving that stay true when the hardware changes: memory math, prefill vs decode, the KV cache, and the latency that matters.
Deploy one model and runtime, run concurrency tests, calculate memory and capacity, measure TTFT and throughput, and produce a build-vs-buy analysis.
- The Memory Math of Serving a Model7 minThe weights fit in VRAM — so why can't you serve 100 long-context users? Because weights are the small, static part of the bill, and the KV cache, which scales with batch size and sequence length, is what actually blows the budget.
- Prefill Versus Decode: Two Phases, Two Bottlenecks7 minProcessing the prompt and generating tokens are different workloads — one compute-bound, one memory-bandwidth-bound. Optimize them as if they were one and you lose both.
- KV-Cache Management and PagedAttention7 minThe KV cache, not raw compute, is the real constraint on how many users you can serve at once. Paging it like operating-system virtual memory is what made high-throughput LLM serving practical.
- The Inference Optimization Toolkit7 minQuantization, speculative decoding, and the parallelism strategies: what each one buys you, what it costs, and when it is the wrong tool. Diagnose the bottleneck first, then pick the one technique that moves it.
- Measuring the Right Latency6 minTokens-per-second alone is a vanity metric. TTFT, inter-token latency, and cost per successful task are what a user and a budget actually feel.
- Managed API Versus Self-Hosting: An Engineering Decision7 minNeither path is intrinsically cheaper or safer. The honest comparison turns on four hinges — utilization, control, contracts, and your operating model — and utilization is where the money crosses over.
- Deploy and Load-Test: The Module Artifact7 minThe capstone: stand up one model on one runtime, push concurrency until it breaks, and turn TTFT and throughput numbers into a defensible build-vs-buy answer.
Emerging Frontiers
Reading the research frontier without getting fooled: inference-time reasoning, state-space models, test-time learning and memory, and self-improving systems.
Compare retrieval, reasoning, and memory approaches on a real task and evaluate each new paradigm critically rather than by its headline.
- Two Ways to Scale: Training-Time Versus Inference-Time Compute7 minFor a decade, "make it better" meant "make it bigger." The newer axis spends compute at request time instead — and it changes the cost curve completely.
- Inference-Time Reasoning Systems7 mino1 and R1 did not get smarter by getting bigger — they spend more compute per question. The durable idea is extra inference-time computation and intermediate reasoning state: sampling, verification, and search, not whether the UI shows the "thinking."
- Beyond Attention: State-Space Models7 minAttention's cost grows with the square of sequence length. State-space models like Mamba swap the quadratic attention matrix for linear-time recurrence — winning on long-sequence throughput but losing on exact recall, which is why the frontier runs hybrids.
- Test-Time Learning and External Memory7 min"Infinite context" is a headline, not a mental model. This lesson replaces it with the honest version: ultra-long-context, external memory, and models that adapt their own state as they read — plus the six limits that survive even an endless stream.
- Self-Improving Systems: A Research Question, Not a Product6 minBreak the phrase apart. Each mechanism people file under "self-improving" — memory, prompt optimization, code modification, fine-tuning from experience — has its own maturity level and its own safety profile, and only some of them are safe to deploy.
- Reading the Frontier Without Getting Fooled7 minThe capstone skill of the whole series: how to evaluate a new-paradigm claim by separating the durable mechanism from the demo, the benchmark from the transfer, and pricing what it actually costs in latency, compute, and attack surface.