The beginner series · 4 min
Prompting vs. RAG vs. Fine-Tuning: Which One Do You Actually Need?
Which lever to pull when an AI system is failing.
Prompting changes the instructions. RAG changes the information the model can see. Fine-tuning changes the model's learned behavior. Pick by the failure you actually have: vague task → prompt it better; missing or fast-changing facts → give it RAG or tools; the same wrong behavior over and over on good inputs → then consider fine-tuning. Most teams reach for fine-tuning first, and it's usually the wrong lever.
Reading promise: No assumed AI knowledge. Jargon gets translated before it gets used.
Meet Steve
Pretend you hired an extremely intelligent employee named Steve. Steve produces a terrible weekly security report. There are several completely different reasons that could happen.
Problem 1: Steve Does Not Understand the Assignment
You asked for “the weekly report” but never specified audience, length, sections, or what matters. You clarify: executive summary, material incidents, metrics, outstanding risks, under two pages, no raw log dumps. Steve immediately improves.
That was an instruction problem. Start with prompting.
Problem 2: Steve Does Not Have the Information
You ask, “How many malware incidents did we have last quarter?” The answer lives in your incident database. Telling Steve to “think deeply” does not create access to that database.
That was an information problem. Use retrieval or tools.
Problem 3: Steve Repeatedly Behaves the Wrong Way
Steve sees the right information and understands the task, but repeatedly violates the same output conventions or decision pattern. You have many examples of what correct behavior looks like.
That may be a learned-behavior problem. Fine-tuning becomes worth considering.
The Mental Model
PROMPTING Tell it what you want. RAG / TOOLS Give it information it needs. FINE-TUNING Train the model toward repeated patterns of behavior.
OpenAI’s current accuracy-optimization guidance treats prompt engineering, retrieval, and fine-tuning as different levers and frames the choice around the type of failure you are trying to fix.[1]
Prompting Changes the Request, Not the Model
You are assisting a security analyst. Classify the authentication sequence as: - benign - suspicious - likely compromise Rules: - do not mark malicious solely due to geography - cite the events supporting the assessment - return no more than five bullets
Same model. Better instructions. You can also add examples—few-shot prompting—inside the context without changing model weights.
RAG Changes What the Model Knows at Runtime
Ask, “What is our policy for sending confidential files to personal email?” A base model may know typical corporate practices. You need your current policy.
This is especially natural for private, rapidly changing, user-specific, or very large information. OpenAI’s model optimization guidance recommends providing relevant external context when the answer depends on information outside the model’s training data.[2]
Fine-Tuning Changes the Model
Fine-tuning continues training on task-specific examples so model parameters move toward the behaviors represented in that dataset. The important word is behavior: patterns of classification, structure, style, terminology, or tool usage—not “use this as a database.”
BASE MODEL │ ├─ training examples: │ input → preferred output │ input → preferred output │ input → preferred output ▼ FINE-TUNED MODEL
Why Fine-Tuning Is Not a Knowledge Base
If the company handbook changes vacation from 20 to 25 days, a retrieval system can index the updated policy. If you tried to treat fine-tuning as document storage, you would be asking a learned statistical model to behave like a versioned records system. That is usually the wrong abstraction.
Precise, current record? → database / retrieval Repeated behavior? → maybe fine-tuning
A Cybersecurity Example Using All Three
Imagine an investigation assistant.
| Layer | What it contributes |
|---|---|
| Prompt | “Never call an incident confirmed malicious without evidence. Return Summary, Evidence, Assessment, Confidence, Next Steps.” |
| RAG/tools | Current SIEM events, EDR telemetry, identity data, internal procedures, recent threat intelligence. |
| Fine-tuning | Recurring analyst classification patterns, preferred severity conventions, consistent structure or specialized terminology. |
These approaches are not mutually exclusive. OpenAI explicitly notes that context problems and behavior problems can coexist and the techniques can be stacked.[1]
The Decision Tree
BAD ANSWER
│
▼
Did the model have the necessary information?
│
┌─┴─┐
NO YES
│ │
▼ ▼
RAG/ Did it clearly understand the desired task/output?
TOOLS │
┌─┴─┐
NO YES
│ │
▼ ▼
PROMPT Does the same behavior keep failing
across representative examples?
│
┌─┴─┐
YES NO
│ │
▼ ▼
CONSIDER keep debugging/evaluating
FINE-TUNEDo Evals Before You Start Adding Machinery
“The model isn’t good enough” is not a diagnosis. Build a representative test set and categorize failures.
100 evaluation cases 72 correct 18 wrong because required information was missing 4 wrong because output format was violated 6 wrong despite apparently correct information Now you have engineering questions instead of vibes.
OpenAI’s optimization workflow puts evaluation and a prompt baseline before fine-tuning because you need to know what is failing—and whether the intervention helped.[2]
The Cheat Sheet
| Problem | First lever |
|---|---|
| The assignment is vague | Prompting |
| The answer requires private/current facts | RAG or tools |
| The same output behavior is inconsistent | Prompt first; consider fine-tuning if the gap persists |
| Company docs change frequently | RAG |
| Need live system state | Tools/API/MCP |
| Do not know why quality is poor | Evals |
Prompting: tell it. RAG: show it. Fine-tuning: train it.
Sources
- [1] OpenAI — Optimizing LLM accuracy. — https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy
- [2] OpenAI — Model optimization. — https://developers.openai.com/api/docs/guides/model-optimization
- [3] OpenAI — Fine-tuning best practices. — https://developers.openai.com/api/docs/guides/fine-tuning-best-practices