Evaluation, Benchmarking & Observability · 7 min

The CI Evaluation Harness: Fail the Build on Regression

The module's capstone: wire your golden, adversarial, security, latency, and cost checks into CI so a quality or safety regression turns the build red and blocks the merge before it reaches users.

A green test suite tells you the code runs. It says nothing about whether your model started refusing valid queries, whether last week's prompt tweak quietly doubled token spend, or whether someone can now jailbreak your support bot by pasting a Base64 blob. None of those failures throw an exception. They pass every unit test you own and ship straight to users. The fix is to treat evaluation the way you treat a failing assertion: wire it into CI, set thresholds, and let a regression turn the build red before it turns into an incident.

This is the capstone for the module. You've built golden sets, adversarial probes, judges, RAG metrics, latency and cost meters. On their own, each is a dashboard someone glances at on Tuesdays. Bolted to a pipeline that can block a merge, they become a contract.

Evaluation only matters if it can block a deploy

Here is the uncomfortable part. A benchmark you look at is advisory. A benchmark that can fail a build is enforcement. The difference is entirely about where it runs and what it is allowed to do.

Think about the "build, break, defend, measure" loop from the security work in Modules 6 and 7. You find an attack that works, you defend against it, then you write a test that reproduces the original break. That test is now permanent. It runs on every commit forever. If a refactor, a model swap, or a prompt edit reopens the hole, the build goes red and the PR can't merge. The attack you already paid to discover never comes back for free. Regression-as-permanent-test is the whole game. A CI eval harness is that idea applied to quality, safety, latency, and cost at once.

The mechanics look like ordinary testing because they should. If your team can already read a pytest failure, they can already read an eval failure. Don't invent a new ceremony.

The shape of the gate

A gate is a script that loads a dataset, runs your system over it, scores the outputs, compares scores to thresholds, and exits non-zero if any threshold is breached. That non-zero exit is the entire point. It is the signal CI reads to mark the job failed.

Structure it like a test file. Here is a golden-set gate written as parametrized pytest cases:

import json
import pytest
from myapp.pipeline import answer
from myapp.judge import score  # returns 0.0-1.0 relevance/correctness

GOLDEN = json.load(open("evals/golden.json"))

# Per-slice floors. A global average hides the slice that fell off a cliff.
THRESHOLDS = {
    "billing":   0.85,
    "technical": 0.80,
    "default":   0.75,
}

@pytest.mark.parametrize("case", GOLDEN, ids=lambda c: c["id"])
def test_golden_quality(case):
    out = answer(case["question"])
    s = score(case["question"], out, case["reference"])
    floor = THRESHOLDS.get(case["slice"], THRESHOLDS["default"])
    assert s >= floor, f"{case['id']} scored {s:.2f} < {floor} ({case['slice']})"

Two choices make this real rather than decorative. First, thresholds are per slice, not one global number. A 0.82 average can hide your billing questions collapsing from 0.90 to 0.60 while everything else drifts up a point. Users don't experience the average; they experience the slice they landed in. Second, the reference answers and the judge come from earlier lessons. You are reusing infrastructure, not standing up a parallel one.

Run that gate against a regressed model and CI hands you exactly the signal you want:

$ pytest evals/ -q
FAILED evals/test_golden.py::test_golden_quality[billing-014]
  billing-014 scored 0.61 < 0.85 (billing)
1 failed, 47 passed in 12.4s
# exit status 1  ->  the job is red, the merge is blocked

The security suite is the same shape with a blunter assertion. Load your jailbreak and prompt-injection corpus, run each case through the system, and fail if any forbidden behavior appears:

ATTACKS = json.load(open("evals/redteam.json"))

@pytest.mark.parametrize("atk", ATTACKS, ids=lambda a: a["id"])
def test_no_policy_violation(atk):
    out = answer(atk["prompt"])
    assert not violates_policy(out), (
        f"{atk['id']} ({atk['technique']}) produced a violation:\n{out[:200]}"
    )

Quality gates are usually thresholds: "at least 80%." Security gates are usually absolutes: zero tolerance for the attacks you have already defended against. A 95% pass rate on your jailbreak suite means one in twenty known attacks works, which is not a pass. Wire that suite so a single regression fails the job.

There are two ways to set the floor, and they solve different problems. An absolute threshold ("billing must clear 0.85") is simple and easy to reason about, but you have to pick the number, and pick it wrong and the gate is either useless or permanently red. A baseline-relative gate compares this run to the last known-good run and fails on a drop larger than some tolerance: assert score >= baseline["billing"] - 0.03. That catches slow erosion an absolute floor sleeps through — a slice sliding from 0.91 to 0.86 to 0.81 over three PRs never trips a 0.80 floor, but each step trips a three-point delta. The cost is that you now store and version a baseline, and refresh it deliberately when you ship a real improvement. Most mature harnesses run both: an absolute floor for the line you never cross, and a delta check for the drift you would otherwise normalize.

Budgets and latency are assertions too

Cost and latency regress silently because nothing crashes. A model swap that improves answer quality by two points and triples the token count is not obviously a win, and no unit test will tell you. Make the budget an assertion.

def test_cost_and_latency_budget():
    result = run_suite(GOLDEN)          # tokens + wall-clock per case
    p95_latency = percentile([r.ms for r in result], 95)
    avg_cost = mean([r.tokens * PRICE_PER_TOKEN for r in result])

    assert p95_latency <= 2500, f"p95 latency {p95_latency}ms > 2500ms budget"
    assert avg_cost <= 0.004,   f"avg cost ${avg_cost:.4f} > $0.004/query budget"

Use p95, not the mean. Mean latency looks fine while your tail grows, and the tail is real people who wait eight seconds and leave. Percentiles are how you catch the tail before it catches you; the Google SRE book treats latency as one of its four golden signals for exactly this reason. Price the run in dollars, not tokens, so the number in the failure message is the number your finance team already tracks.

Wiring it into the pipeline

The harness runs like any other job. A minimal GitHub Actions step:

- name: Evaluation gate
  run: pytest evals/ -v --tb=short
  # non-zero exit fails the job and blocks the merge

The job going red is necessary but not sufficient — a red check that nobody is required to pass is just a suggestion with extra steps. On GitHub, mark the eval job a required status check in the branch's protection rules so the merge button actually greys out until it passes. That one setting is what converts your assertions from advice into a gate. Without it, the loop from the top of this lesson is incomplete: you can measure the regression and still ship it.

A few more decisions separate a harness people trust from one they route around.

Keep it fast and deterministic where you can. A gate that takes twenty minutes and flakes twice a week gets disabled within a month. Pin decoding to temperature=0 for scored runs, cache model responses per input hash so unchanged prompts don't re-bill, and keep the CI dataset to the tens or low hundreds of cases that actually discriminate between a good build and a bad one. Save the ten-thousand-case sweep for a nightly job, not the per-PR gate.

Handle judge noise honestly. If an LLM grades your outputs, its own variance can flip a borderline case and fail a build for no real reason. temperature=0 narrows that variance but does not erase it, since batching and hardware can still shift a token. Two defenses hold up: set thresholds with a margin below observed-stable performance so normal jitter doesn't trip them, and for cases that must be exact, prefer a cheap deterministic check — a regex, a JSON-schema validation, an exact match — over a model judge. The field has not settled on how to make LLM-judge scores fully reproducible, so don't treat a single graded number as ground truth. Gate on the metrics you can defend.

Separate a hard fail from a warning. A security violation blocks. A one-point quality wobble might annotate the PR instead. Tools like promptfoo and OpenAI Evals exist to express "run these cases, score them, assert these thresholds" without hand-rolling the plumbing; conceptually they do what the pytest sketches above do. Use one if it fits, but understand the mechanism so you can tune what fails versus what merely complains.

What "green" should mean

When you are done, a green pipeline should assert more than "it compiles." It should mean quality held on every slice, no known attack got through, p95 latency stayed under budget, and average cost per query didn't creep. That is a claim worth making on every merge.

Start smaller than you think. Ten golden cases, your top five defended attacks, one latency assertion, one cost assertion. Get that failing correctly — deliberately break something and watch the build go red — before you expand. A harness you don't trust to fail honestly is worse than none, because people learn to ignore it. The one you trust is the one that catches the regression at 2pm on a Tuesday instead of in a postmortem.

Sources