The beginner series · 4 min

Prompting vs. RAG vs. Fine-Tuning: Which One Do You Actually Need?

Which lever to pull when an AI system is failing.

Prompting changes the instructions. RAG changes the information the model can see. Fine-tuning changes the model's learned behavior. Pick by the failure you actually have: vague task → prompt it better; missing or fast-changing facts → give it RAG or tools; the same wrong behavior over and over on good inputs → then consider fine-tuning. Most teams reach for fine-tuning first, and it's usually the wrong lever.

Reading promise: No assumed AI knowledge. Jargon gets translated before it gets used.

Meet Steve

Pretend you hired an extremely intelligent employee named Steve. Steve produces a terrible weekly security report. There are several completely different reasons that could happen.

Problem 1: Steve Does Not Understand the Assignment

You asked for “the weekly report” but never specified audience, length, sections, or what matters. You clarify: executive summary, material incidents, metrics, outstanding risks, under two pages, no raw log dumps. Steve immediately improves.

That was an instruction problem. Start with prompting.

Problem 2: Steve Does Not Have the Information

You ask, “How many malware incidents did we have last quarter?” The answer lives in your incident database. Telling Steve to “think deeply” does not create access to that database.

That was an information problem. Use retrieval or tools.

Problem 3: Steve Repeatedly Behaves the Wrong Way

Steve sees the right information and understands the task, but repeatedly violates the same output conventions or decision pattern. You have many examples of what correct behavior looks like.

That may be a learned-behavior problem. Fine-tuning becomes worth considering.

The Mental Model

PROMPTING
Tell it what you want.

RAG / TOOLS
Give it information it needs.

FINE-TUNING
Train the model toward repeated patterns of behavior.

OpenAI’s current accuracy-optimization guidance treats prompt engineering, retrieval, and fine-tuning as different levers and frames the choice around the type of failure you are trying to fix.[1]

Prompting Changes the Request, Not the Model

You are assisting a security analyst.

Classify the authentication sequence as:
- benign
- suspicious
- likely compromise

Rules:
- do not mark malicious solely due to geography
- cite the events supporting the assessment
- return no more than five bullets

Same model. Better instructions. You can also add examples—few-shot prompting—inside the context without changing model weights.

RAG Changes What the Model Knows at Runtime

Ask, “What is our policy for sending confidential files to personal email?” A base model may know typical corporate practices. You need your current policy.

QUESTION
search current policy library
retrieve exact DLP policy section
put it in context
LLM answers with source

This is especially natural for private, rapidly changing, user-specific, or very large information. OpenAI’s model optimization guidance recommends providing relevant external context when the answer depends on information outside the model’s training data.[2]

Fine-Tuning Changes the Model

Fine-tuning continues training on task-specific examples so model parameters move toward the behaviors represented in that dataset. The important word is behavior: patterns of classification, structure, style, terminology, or tool usage—not “use this as a database.”

BASE MODEL
   │
   ├─ training examples:
   │    input → preferred output
   │    input → preferred output
   │    input → preferred output
   ▼
FINE-TUNED MODEL

Why Fine-Tuning Is Not a Knowledge Base

If the company handbook changes vacation from 20 to 25 days, a retrieval system can index the updated policy. If you tried to treat fine-tuning as document storage, you would be asking a learned statistical model to behave like a versioned records system. That is usually the wrong abstraction.

Precise, current record? → database / retrieval
Repeated behavior?       → maybe fine-tuning

A Cybersecurity Example Using All Three

Imagine an investigation assistant.

LayerWhat it contributes
Prompt“Never call an incident confirmed malicious without evidence. Return Summary, Evidence, Assessment, Confidence, Next Steps.”
RAG/toolsCurrent SIEM events, EDR telemetry, identity data, internal procedures, recent threat intelligence.
Fine-tuningRecurring analyst classification patterns, preferred severity conventions, consistent structure or specialized terminology.

These approaches are not mutually exclusive. OpenAI explicitly notes that context problems and behavior problems can coexist and the techniques can be stacked.[1]

The Decision Tree

BAD ANSWER
   │
   ▼
Did the model have the necessary information?
   │
 ┌─┴─┐
NO  YES
│    │
▼    ▼
RAG/ Did it clearly understand the desired task/output?
TOOLS   │
      ┌─┴─┐
     NO  YES
     │    │
     ▼    ▼
   PROMPT Does the same behavior keep failing
          across representative examples?
             │
           ┌─┴─┐
          YES  NO
          │    │
          ▼    ▼
       CONSIDER keep debugging/evaluating
      FINE-TUNE

Do Evals Before You Start Adding Machinery

“The model isn’t good enough” is not a diagnosis. Build a representative test set and categorize failures.

100 evaluation cases

72 correct
18 wrong because required information was missing
 4 wrong because output format was violated
 6 wrong despite apparently correct information

Now you have engineering questions instead of vibes.

OpenAI’s optimization workflow puts evaluation and a prompt baseline before fine-tuning because you need to know what is failing—and whether the intervention helped.[2]

The Cheat Sheet

ProblemFirst lever
The assignment is vaguePrompting
The answer requires private/current factsRAG or tools
The same output behavior is inconsistentPrompt first; consider fine-tuning if the gap persists
Company docs change frequentlyRAG
Need live system stateTools/API/MCP
Do not know why quality is poorEvals
Prompting: tell it. RAG: show it. Fine-tuning: train it.

Sources

Prompting vs. RAG vs. Fine-Tuning: Which One Do You Actually Need? — WTF Is…? — Plain-English AI · AdversariaLLM