The beginner series · 4 min
WTF Is an AI Agent? And How Is That Different From a Chatbot?
How models move from answering to taking iterative action.
The difference is not a smarter model. A chatbot answers and stops. An agent is handed a goal and then loops — decide the next step, use a tool, look at the result, decide again — until it's done. Same kind of model underneath; what's new is the loop and the hands (tools) it can reach for. That's also why an agent raises the stakes: a chatbot can be wrong, but an agent can be wrong and then act on it.
Reading promise: No assumed AI knowledge. Jargon gets translated before it gets used.
Chatbot vs. Agent in One Picture
CHATBOT
You ask → LLM generates answer → done
AGENT
Goal → LLM decides next action → tool executes → result
↑ │
└──────── reason again ────────────┘
↓
doneOpenAI describes agent systems around models, tools, instructions, and a loop that continues until an exit condition. Anthropic makes a useful distinction between workflows, where code predetermines the path, and agents, where the model dynamically directs its process and tool use.[1][2]
Give the Brain Hands
A language model can decide, “I should check the calendar.” That thought does not inspect a calendar. The application needs tools.
LLM
│
┌─────────┼─────────┐
▼ ▼ ▼
Calendar Web Email
tool search toolTools can be APIs, database calls, shell commands, browser controls, code execution, MCP capabilities, or specialized application functions.
Tools Alone Do Not Automatically Make an Agent
Consider two systems.
FIXED WORKFLOW
1. Search database
2. Summarize result
3. Email summary
The programmer chose the sequence.
AGENT
Goal → model asks "what do I need?"
→ chooses tool
→ sees result
→ decides if another step is required
→ repeatsBoth can be valuable. “Agentic” is not automatically better. If the path is known and deterministic, a workflow may be cheaper, faster, and easier to audit.
The Agent Loop
┌──────────┐
│ GOAL │
└────┬─────┘
▼
┌──────────┐
│ LLM │
│ What now?│
└────┬─────┘
┌───────┴────────┐
▼ ▼
final answer? tool call
│ │
yes ▼
│ execute tool
▼ │
DONE observe result
│
└────→ LLM againOpenAI’s technical description of the Codex harness shows this directly: the model either returns a final response or asks for a tool; the tool result is put back into context; the model is invoked again.[3]
A Website-Debugging Example
You say: “Figure out why the website returns HTTP 500.” A chatbot can list generic causes. An agent with safe access to logs and configuration can investigate.
Goal: explain HTTP 500s ↓ read_logs() → DB connection timeout ↓ test_connectivity() → network path works ↓ inspect_config() → hostname points at retired DB ↓ Conclusion: likely stale database endpoint
The path emerged from observations. Nobody had to hard-code “logs, then network test, then config” for this exact incident.
Cybersecurity Is Almost a Perfect Agent Example
Security investigations naturally branch based on evidence. Start with an impossible-travel alert:
ALERT: Alice logged in from Germany ↓ Check sign-in history → New York 22 minutes earlier ↓ Check German IP → commercial VPN ↓ Check Alice's endpoint → corporate laptop active in New York ↓ Check who else used same IP → six employees ↓ Check corporate VPN egress → same provider ↓ Assessment: likely VPN artifact, not compromise
Each observation changes the next useful question. That is exactly where dynamic tool selection can help.
What Is an Agent Harness?
The LLM is not usually responsible for every operational detail. Software around it handles tool execution, permissions, state, retries, logs, timeouts, approvals, context, sandboxes, and stopping conditions. That surrounding runtime is often called an agent harness.
┌──────────────── AGENT HARNESS ────────────────┐ │ instructions │ │ ↓ │ │ LLM ←→ context / state │ │ │ │ │ ├─ tools │ │ ├─ approvals │ │ ├─ sandbox │ │ ├─ logging │ │ └─ stopping conditions │ └───────────────────────────────────────────────┘
Single-Agent vs. Multi-Agent
Multi-agent systems give different agents specialized responsibilities or let a manager delegate work. OpenAI describes patterns such as a central manager calling specialist agents or agents handing work to one another.[1]
MANAGER
┌───────┼───────┐
▼ ▼ ▼
RESEARCH ANALYZE WRITE
└───────┼───────┘
▼
REVIEWMore agents are not automatically more intelligent. They can also mean more latency, more tokens, more coordination errors, more cost, and more places to fail.
Autonomy Is a Spectrum
LOW AUTONOMY model suggests an action read-only tool calls multiple investigation steps writes after human approval automatic writes long-running independent work HIGH AUTONOMY
High-impact actions do not need to be fully autonomous. A good security agent might investigate independently, then stop before isolate_host() or disable_user() and ask for approval.
Why Agent Errors Can Be More Dangerous
CHATBOT ERROR
"I think the account is compromised."
AGENT ERROR
"I think the account is compromised."
↓
disable_user()
↓
CEO is locked outThe ability to act is what makes agents useful—and what raises the stakes. Guardrails should be layered: least privilege, tool allowlists, approvals, budgets, sandboxes, maximum steps, audit logs, and independent validation where appropriate.
The Cheat Sheet
| System | What defines it |
|---|---|
| Chatbot / one-shot LLM | Input → model → output |
| Workflow | A programmer defines the path |
| Agent | The model dynamically chooses actions based on observations |
| Multi-agent system | Several agents coordinate or specialize |
| Agent harness | Software that runs the loop and controls tools, state, permissions, and context |
A useful beginner formula: LLM + instructions + tools + a decision/action/observation loop ≈ agent.
Sources
- [1] OpenAI — A practical guide to building agents. — https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf
- [2] Anthropic — Building Effective AI Agents. — https://www.anthropic.com/engineering/building-effective-agents
- [3] OpenAI — Unrolling the Codex agent loop.