The beginner series · 4 min

WTF Is an AI Agent? And How Is That Different From a Chatbot?

How models move from answering to taking iterative action.

The difference is not a smarter model. A chatbot answers and stops. An agent is handed a goal and then loops — decide the next step, use a tool, look at the result, decide again — until it's done. Same kind of model underneath; what's new is the loop and the hands (tools) it can reach for. That's also why an agent raises the stakes: a chatbot can be wrong, but an agent can be wrong and then act on it.

Reading promise: No assumed AI knowledge. Jargon gets translated before it gets used.

Chatbot vs. Agent in One Picture

CHATBOT
You ask → LLM generates answer → done

AGENT
Goal → LLM decides next action → tool executes → result
        ↑                                  │
        └──────── reason again ────────────┘
                         ↓
                       done

OpenAI describes agent systems around models, tools, instructions, and a loop that continues until an exit condition. Anthropic makes a useful distinction between workflows, where code predetermines the path, and agents, where the model dynamically directs its process and tool use.[1][2]

Give the Brain Hands

A language model can decide, “I should check the calendar.” That thought does not inspect a calendar. The application needs tools.

                 LLM
                  │
        ┌─────────┼─────────┐
        ▼         ▼         ▼
     Calendar    Web       Email
       tool     search      tool

Tools can be APIs, database calls, shell commands, browser controls, code execution, MCP capabilities, or specialized application functions.

Tools Alone Do Not Automatically Make an Agent

Consider two systems.

FIXED WORKFLOW
1. Search database
2. Summarize result
3. Email summary

The programmer chose the sequence.

AGENT
Goal → model asks "what do I need?"
     → chooses tool
     → sees result
     → decides if another step is required
     → repeats

Both can be valuable. “Agentic” is not automatically better. If the path is known and deterministic, a workflow may be cheaper, faster, and easier to audit.

The Agent Loop

             ┌──────────┐
             │   GOAL   │
             └────┬─────┘
                  ▼
             ┌──────────┐
             │   LLM    │
             │ What now?│
             └────┬─────┘
          ┌───────┴────────┐
          ▼                ▼
     final answer?      tool call
          │                │
         yes               ▼
          │           execute tool
          ▼                │
        DONE          observe result
                           │
                           └────→ LLM again

OpenAI’s technical description of the Codex harness shows this directly: the model either returns a final response or asks for a tool; the tool result is put back into context; the model is invoked again.[3]

A Website-Debugging Example

You say: “Figure out why the website returns HTTP 500.” A chatbot can list generic causes. An agent with safe access to logs and configuration can investigate.

Goal: explain HTTP 500s
  ↓
read_logs()
  → DB connection timeout
  ↓
test_connectivity()
  → network path works
  ↓
inspect_config()
  → hostname points at retired DB
  ↓
Conclusion: likely stale database endpoint

The path emerged from observations. Nobody had to hard-code “logs, then network test, then config” for this exact incident.

Cybersecurity Is Almost a Perfect Agent Example

Security investigations naturally branch based on evidence. Start with an impossible-travel alert:

ALERT: Alice logged in from Germany
  ↓
Check sign-in history
  → New York 22 minutes earlier
  ↓
Check German IP
  → commercial VPN
  ↓
Check Alice's endpoint
  → corporate laptop active in New York
  ↓
Check who else used same IP
  → six employees
  ↓
Check corporate VPN egress
  → same provider
  ↓
Assessment: likely VPN artifact, not compromise

Each observation changes the next useful question. That is exactly where dynamic tool selection can help.

What Is an Agent Harness?

The LLM is not usually responsible for every operational detail. Software around it handles tool execution, permissions, state, retries, logs, timeouts, approvals, context, sandboxes, and stopping conditions. That surrounding runtime is often called an agent harness.

┌──────────────── AGENT HARNESS ────────────────┐
│ instructions                                  │
│      ↓                                        │
│     LLM ←→ context / state                    │
│      │                                        │
│      ├─ tools                                 │
│      ├─ approvals                             │
│      ├─ sandbox                               │
│      ├─ logging                               │
│      └─ stopping conditions                   │
└───────────────────────────────────────────────┘

Single-Agent vs. Multi-Agent

Multi-agent systems give different agents specialized responsibilities or let a manager delegate work. OpenAI describes patterns such as a central manager calling specialist agents or agents handing work to one another.[1]

                    MANAGER
              ┌───────┼───────┐
              ▼       ▼       ▼
          RESEARCH  ANALYZE  WRITE
              └───────┼───────┘
                      ▼
                   REVIEW
More agents are not automatically more intelligent. They can also mean more latency, more tokens, more coordination errors, more cost, and more places to fail.

Autonomy Is a Spectrum

LOW AUTONOMY
  model suggests an action
  read-only tool calls
  multiple investigation steps
  writes after human approval
  automatic writes
  long-running independent work
HIGH AUTONOMY

High-impact actions do not need to be fully autonomous. A good security agent might investigate independently, then stop before isolate_host() or disable_user() and ask for approval.

Why Agent Errors Can Be More Dangerous

CHATBOT ERROR
"I think the account is compromised."

AGENT ERROR
"I think the account is compromised."
      ↓
disable_user()
      ↓
CEO is locked out

The ability to act is what makes agents useful—and what raises the stakes. Guardrails should be layered: least privilege, tool allowlists, approvals, budgets, sandboxes, maximum steps, audit logs, and independent validation where appropriate.

The Cheat Sheet

SystemWhat defines it
Chatbot / one-shot LLMInput → model → output
WorkflowA programmer defines the path
AgentThe model dynamically chooses actions based on observations
Multi-agent systemSeveral agents coordinate or specialize
Agent harnessSoftware that runs the loop and controls tools, state, permissions, and context
A useful beginner formula: LLM + instructions + tools + a decision/action/observation loop ≈ agent.

Sources

  • [3] OpenAI — Unrolling the Codex agent loop.