Evaluation, Benchmarking & Observability · 7 min

Reading a Benchmark Without Getting Fooled

MMLU and SWE-bench measure two different philosophies of "smart." What each actually tests, why training contamination quietly inflates scores, and five questions that tell you whether a number transfers to your problem.

Two numbers get quoted the same way in press releases: "92% on MMLU" and "43% on SWE-bench." They look like the same kind of fact — a model took a test, here's the grade. They are not the same kind of fact at all. One is a multiple-choice exam where the model picks A, B, C, or D across 57 subjects. The other drops the model into a real Python repository with a real bug report and grades it on whether the patch it writes makes a failing test suite pass. A model can be brilliant at the first and useless at the second. If you read both scores as "how smart is this model," you will buy the wrong model.

This lesson is about reading a benchmark the way an adversary reads a contract: looking for what it actually binds, not what the headline implies.

Two benchmarks, two philosophies

MMLU (Measuring Massive Multitask Language Understanding) is a knowledge exam. Hendrycks and colleagues assembled about 16,000 multiple-choice questions across 57 subjects — everything from elementary math to US history to professional law and clinical medicine. Each question has four options and exactly one right answer. Scoring is trivial: did the model's chosen letter match the key? A random guesser scores 25%. The whole point of the design is breadth — it asks whether a model has absorbed a wide cross-section of human knowledge.

MMLU item (paraphrased shape):
  Q: One reason governments discourage monopolies is that...
     (A) ...   (B) ...   (C) ...   (D) ...
  Model outputs a single letter. Grader checks: letter == key.

SWE-bench (Jimenez et al.) is the opposite philosophy. It takes 2,294 real issues and pull requests from twelve popular Python projects — Django, scikit-learn, matplotlib, and others. The model gets the issue text and the actual repository at the commit right before the fix. It has to produce a code patch. The harness applies that patch and runs the project's own test suite. You pass only if the tests that were failing now pass and the tests that were already passing still pass.

SWE-bench task:
  input:  issue text + full repo @ base commit
  output: a git diff (patch)
  grade:  apply patch -> run FAIL_TO_PASS + PASS_TO_PASS tests
          score = 1 only if all target tests go green

When SWE-bench launched in 2023, the best model in the paper resolved under 2% of issues; the climb into the 40s came largely from agent scaffolding wrapped around the model, not from the base model suddenly getting smarter. That alone is a reason to ask what a headline number is really counting.

Nothing about SWE-bench is multiple choice. There is no answer key to match against — there is a runtime that either goes green or doesn't. That single difference cascades into everything else you care about.

The five questions to ask any benchmark

Whenever someone waves a score at you, run it through these five. They take about two minutes and they save you from most benchmark theater.

1. What capability is actually measured? MMLU measures recall and reasoning over static knowledge. SWE-bench measures the ability to localize a bug in unfamiliar code, edit multiple files, and produce something that runs. "GPT-4 scores well on MMLU" tells you it knows a lot of facts. It tells you almost nothing about whether it can fix your Django app. Name the capability in plain words before you trust the number.

2. What is the interaction mode — pick, generate, execute, or agent? This is the axis that MMLU and SWE-bench split on, and it predicts how the score behaves:

  • Pick (MMLU): the model selects from given options. High ceiling for guessing, easy to grade, easy to game.
  • Generate: the model writes free text graded by string match or a model judge, as in summarization or translation.
  • Execute (SWE-bench): the output is run against ground truth, so it is hard to fake — code either works or it doesn't.
  • Agent: the model takes multiple actions against an environment over many turns, reading output and deciding again.

The further down this list, the more the benchmark resembles real work, and the harder it is to inflate the score without actually having the capability. Execution grading is the reason a 43% on SWE-bench is a genuinely impressive number and a 92% on a multiple-choice test is more suspicious than it looks.

3. Could training contamination explain the score? This is the big one. A benchmark is only valid if the model hasn't already seen the answers during training. MMLU questions and their answer keys have sat on the public web for years — in the original repo, in scraped copies, in blog posts, in dataset mirrors. If those pages landed in a model's training corpus, a high MMLU score partly measures memorization, not understanding. Researchers have measured this gap directly: rephrase or freshly rewrite the questions in the same format, and contaminated models often drop several points. When the drop is large, the original number was scoring the training set, not the skill. This isn't a fringe worry; it's a documented, load-bearing threat to benchmark validity.

SWE-bench has the same exposure in principle — the fixing commits are public on GitHub — which is exactly why the maintainers released SWE-bench Verified, a 500-task human-filtered subset, and why researchers watch for models that "solve" an issue by reproducing the known patch rather than reasoning to it. Ask: when was this benchmark published, is it public, and was the model trained after that date? If yes, yes, and yes, discount the score.

# Contamination smell test: does the model complete a held-out
# benchmark item from a fragment it should not be able to finish?
prompt = benchmark_question[:60]   # first 60 chars only
out = model.generate(prompt, max_tokens=200)
# If it reproduces the exact remaining question text AND the
# right answer verbatim, that item was likely in training.

It's not proof, but verbatim regurgitation of a "held-out" test item is a strong tell.

4. Is scoring deterministic? MMLU is deterministic — letter equals key, no ambiguity. SWE-bench is deterministic in its grader, since tests pass or fail, but noisy in the model's path to it, because generation involves sampling. Benchmarks graded by another LLM ("LLM-as-judge") are the least deterministic: run the same eval twice and the number moves. The same model at temperature 0.8 can swing several points between runs on a hard generation task; report a single seed and you have hidden that swing. If two models sit at 71.2 and 71.5 on a judge-graded benchmark, that gap is probably noise. On MMLU it might be real. Always ask whether the leaderboard reports variance or a single run.

5. Does success transfer to your application? The only question that ultimately matters. Transfer depends on how well the benchmark's task distribution matches yours. SWE-bench is Python, library-scale, test-covered, English issue reports. If your codebase is a 2-million-line C++ monolith with no test suite, a great SWE-bench score is weak evidence — and two models a few points apart there can flip order entirely on your stack. MMLU's law and medicine subjects do not certify a model for legal or clinical use — the format of four options and one clean answer erases the ambiguity that defines real professional work. The honest move is to build a small private eval from your own tasks — even 50 hand-graded examples — and treat public benchmarks as a prior, not a verdict.

Putting it to work

Say a vendor pitches you a coding assistant with "state-of-the-art SWE-bench." Run the five questions. Capability: edits real repos, good. Mode: execution-graded, good, hard to fake. Contamination: is it SWE-bench Verified, and was the model's cutoff before or after the tasks went public? Ask. Determinism: how many runs, and what is the pass@1 versus pass@k — because "resolves 50%" at pass@10 means ten attempts, not one. Transfer: Python and tested; is that you?

The rank on a leaderboard is the least useful number in the room; the shape of the task underneath it is everything. Now the score is a piece of evidence you can weigh, not a slogan you have to swallow. The skill isn't memorizing which model tops which leaderboard this month. It's knowing that "92% on a multiple-choice test" and "43% at fixing real bugs under execution grading" are answers to completely different questions — and being able to say which question you actually need answered.

Sources