Evaluation, Benchmarking & Observability · 7 min

The Evaluation Hierarchy: Five Layers of "Is It Good?"

A benchmark score tells you almost nothing about whether your app works. This lays out the ladder that does — from deterministic component tests up through model, RAG, and agent eval to the business metrics that are the only real verdict.

A model can score 86% on MMLU and still tell your users the wrong thing about your own product on the first try. That gap — between a number on a leaderboard and whether the thing works — is the entire subject of this module, and it's where most teams burn the most time.

Here's the trap. You pick a model because it's near the top of a public benchmark. You wire it into a support assistant over your docs. It ships. Then a customer asks "how do I cancel and get a refund?" and the assistant confidently explains a refund flow you retired eighteen months ago. The model didn't fail a knowledge test. It failed your task, using your data, in your UI, under your latency budget. MMLU never measured any of that, and it was never going to.

The fix isn't a better benchmark. It's understanding that "is it good?" is not one question. It's five, stacked, and each layer measures something the layer beneath it structurally cannot see.

The five layers

Think of evaluation as a ladder. Lower rungs are cheaper, faster, and more objective. Higher rungs are slower, noisier, and closer to the truth you actually care about.

5. Business / task success   "Did the user get what they came for?"
4. Agent / workflow eval     "Did the multi-step process reach a good end state?"
3. Application / RAG eval     "Given our data + prompt, is this answer right?"
2. Model eval                "Is the base model capable in general?"
1. Component tests           "Do the non-LLM parts behave deterministically?"

Read it bottom-up and a rule falls out: a pass at one layer tells you nothing about the layer above. A perfect retriever (layer 1) can feed a perfect prompt to a strong model (layer 2) and still produce a wrong answer (layer 3), because the right chunk was never in the index. A correct answer (layer 3) can sit inside an agent that loops forever and never books the meeting (layer 4). And an agent that books the meeting flawlessly can still miss the point, because the user actually wanted to cancel it (layer 5). Each layer is a different failure surface. Walk them.

Layer 1 and 2: the parts and the engine

Layer 1 is boring in the best way. It's the deterministic scaffolding around the model: does your chunker split on the right boundaries, does your retriever return the documents you expect for a known query, does your JSON parser survive a trailing comma, does your rate limiter back off.

def test_retrieval_returns_current_policy():
    hits = retrieve("refund policy")
    assert hits[0].id == "doc_842"        # the current doc, not the 2024 one
    assert "doc_311" not in [h.id for h in hits]

No model involved, no randomness, run it on every commit. If layer 1 is flaky, every measurement above it is polluted and you won't know why.

Layer 2 is the model in the abstract — the thing benchmarks measure. MMLU is the emblem: 57 subjects, multiple choice, from elementary math to professional law, scored as raw accuracy. It's a genuinely useful signal for general capability. A model at 45% and a model at 85% are different animals, and you should know which you're holding.

But understand exactly what a model benchmark is: a fixed, public, multiple-choice test of a bare model, with no system prompt, no retrieval, no tools, and no user. That's its strength (comparable, reproducible) and its ceiling. Two cracks make it worse than it looks. Contamination: popular benchmarks leak into training corpora, so scores drift upward without capability moving. Construct mismatch: MMLU rewards recalling a fact and picking one of four letters; your app rewards synthesizing three retrieved paragraphs into something a stressed human can act on. Those are not the same skill. This is why leaderboard rank does not equal task success, and why efforts like HELM argue for reporting many metrics across many scenarios instead of one headline number. Use layer 2 to shortlist models. Never to ship one.

Layer 3: your data, your prompt, the first place it gets real

This is where most teams should spend most of their effort, and where the fewest actually do. Layer 3 asks: given our actual system prompt, our actual retrieved context, and a realistic question, is the answer correct, grounded, and complete?

The refund bot lives or dies here. The model was fine. Retrieval returned a stale document, and nothing checked whether the answer was faithful to a current source. RAG evaluation names these failure modes directly. Frameworks like RAGAS split the question into measurable pieces:

context_precision   were the retrieved chunks actually relevant?
context_recall      did we retrieve everything needed to answer?
faithfulness        is every claim in the answer supported by the context?
answer_relevancy    does the answer address the question asked?

Faithfulness is the one that would have caught the refund bug: the answer asserted a refund flow, no retrieved chunk supported it, and that mismatch gets flagged before a customer ever sees it. You build a small dataset of real questions with known-good answers — start with 50, because a curated 50 beats an aspirational 5,000 you never label — run it on every prompt or retrieval change, and watch the numbers move. Often an LLM does the scoring, judging whether each claim is supported. That is its own subject, with its own biases, and it gets a full lesson later in this module.

Layer 4 and 5: the process and the point

Layer 4 shows up the moment the system takes more than one step: it retrieves, calls a tool, reads the result, decides what to do next. Now correctness is about the trajectory, not a single response. Did the agent pick the right tool? Did it recover when the API returned a 500, or loop until it burned the token budget? A travel agent that books a flight to the wrong airport failed at layer 4 even if every individual model call was locally reasonable. You evaluate these by asserting on the end state — was a booking created with the right parameters? — and on the path: how many steps, which tools, where it stalled. That is why traces matter, and why the OpenTelemetry GenAI semantic conventions exist: to record spans for model calls, tool calls, and token counts in a vendor-neutral shape. Instrument for layer 4 now, before you formally evaluate it. You cannot debug a trajectory you didn't record.

Layer 5 is the only one the user actually cares about. Did they resolve the issue? Did they buy the thing, finish onboarding, stop contacting support? These are business metrics — deflection rate, task completion, CSAT, conversion — measured in production, on real traffic, usually behind an A/B test. They are the ground truth. They are also lagging, noisy, and confounded by ten things that have nothing to do with your model. That is precisely why you need layers 1 through 4: they are the fast, controllable proxies that let you improve the system between the slow, honest verdicts layer 5 hands down. The whole point of the ladder is to catch a regression at the cheapest layer that can see it. A broken chunker should fail a layer-1 unit test in seconds, not surface as a two-point dip in CSAT that a data scientist spends three weeks tracing back to a config change nobody remembers.

What to do with this

Don't build all five at once. Build downward from where the pain is. If users complain about wrong answers, you have a layer 3 problem — build a 50-example RAG eval this week. If your agent is erratic, you have a layer 4 problem — add tracing first, then trajectory assertions. Treat the model benchmark as a filter for your shortlist and nothing more.

The rest of this module climbs the ladder: how model benchmarks actually work and where they lie, how to build application and RAG evals, how to make an LLM judge you can trust, how to read agent traces, and how to close the loop with production observability. Carry one belief through all of it: a score is only as meaningful as the layer it was measured at, and the layer that matters most is the one the benchmark never touched.

Sources