The tooling around security LLMs: what orchestrates, what attacks, and what actually measures

D. Rose · 31 July 2026 · Updated 31 August 2026 · 8 min

The frameworks, attack algorithms, and benchmarks in this space are not security models — they run against whatever weights you point them at. A cited tour of the pentest agents, red-team scanners, jailbreak algorithms, and the evaluations that produce real numbers (Cybench, CTIBench, CyberSOCEval, DefenderBench) — and why "has a framework" is not "has reproducible numbers."

There is a category error that runs through this whole field, and it is worth naming before anything else: a piece of tooling is not a model. A framework that drives GPT-4o through a penetration test is not a security-tuned weight. An algorithm that mutates a prompt until a target model complies is not a model either. And a benchmark is not a capability — it is the instrument that turns any of the above into a number you can actually check.

Four verbs keep the categories straight. Some tools orchestrate a model. Some attack or scan one. Some generate adversarial inputs. And some measure. None of them are the security-tuned weights covered in the 2026 survey — they are the scaffolding around whatever model you supply. Conflating the two is how a project with a good README gets mistaken for a model with a good score.

Agents that orchestrate a model

PentestGPT is the reference example: an open-source framework that orchestrates an underlying LLM through a penetration test rather than being a security model itself. The point of naming it here is precisely that it is not a fine-tuned security weight — it is tooling that steers a base model. The same holds for its relatives; the field lists METATRON in this group, positioned as fully offline/local (a project/press claim, not independently verified).

CAI and HackSynth sit in the same category, and both are better documented than a first pass suggests. CAI ("An Open, Bug Bounty-Ready Cybersecurity AI", Mayoral-Vilches et al., Alias Robotics) is an agent framework of specialised sub-agents, co-funded under the EU EIC accelerator. HackSynth (Muzsai, Imolai and Lukács) is a Planner/Summarizer agent for autonomous penetration testing that ships its own evaluation: two CTF benchmark sets built on PicoCTF and OverTheWire, two hundred challenges, with per-model results.

That last point matters for the argument of this post. HackSynth is orchestration and measurement — it is one of the few things in this section that produces a number you can check rather than a capability you have to take on faith.

Tools that attack or scan a model

Garak (NVIDIA) is an LLM vulnerability scanner — it probes a target model for failure modes. PyRIT (Microsoft) is a risk-identification toolkit for generative AI, used to automate red-team probing. Both are model-agnostic: you point them at a model and they test its resilience. Neither ships weights, and neither tells you a model is good — only where it breaks.

CyberSecEval / Purple Llama (Meta) sits on the boundary between tooling and benchmark: a suite that covers insecure-code generation, prompt injection, and offensive/compliance risk. It is worth flagging how it gets misused. One community offensive build in this catalog, BugTraceAI-Apex-G4-26B, reports a self-measured "100% compliance / 0% refusal" on a CyberSecEval MITRE ATT&CK set. That figure measures whether the model answers, not whether its output is correct or functional — a willingness metric wearing a capability metric's clothes. The suite is legitimate; the reading of it there is not.

The attack algorithms

These are methods, not models — recipes for producing adversarial prompts that a defender uses to stress-test resilience.

  • GCG (the llm-attacks codebase) produces adversarial suffixes appended to a prompt. It is the algorithm behind released weight-based generators such as AmpleGCG, OSU NLP's gated Llama-2-based suffix generators — a case of an algorithm hardening into a downloadable (gated) tool.
  • PAIR (Prompt Automatic Iterative Refinement; Chao et al., later at IEEE SaTML 2025) uses one LLM as the attacker against another, refining a semantic prompt across a handful of turns — it advertises jailbreaks in around twenty queries, black-box.
  • TAP (Tree of Attacks with Pruning; Mehrotra et al.) generalises that into a tree search: branch the attacker's candidate prompts, prune the ones judged off-target, and keep the promising line.
  • AutoDAN (Liu, Xu, Chen and Xiao) attacks the stealth problem instead. GCG's suffixes are semantically meaningless and therefore easy to catch on perplexity alone; AutoDAN evolves readable, DAN-style prompts hierarchically, at the sentence and word level, so the result reads like language.
  • Rainbow Teaming (Samvelyan et al., Meta) is not chasing one jailbreak at all. It treats adversarial prompting as a quality-diversity search and produces a spread of effective prompts across a descriptor space — which is what you want when the question is "where does this model break", not "can it be broken".

The through-line is worth having: GCG searches token suffixes, PAIR and TAP search semantic prompts with an LLM in the loop, AutoDAN optimises for readability, and Rainbow Teaming optimises for coverage. Four different answers to what an adversarial prompt is for.

An earlier version of this post said these four appeared "with no individual paper or benchmark attached". That was wrong — every one of them is a published method, and the mechanism notes above are drawn from those papers. Corrected 30 August 2026.

The distinction to carry: an attack algorithm is code you run, not a model you download. Which is the cleanest bridge to the exclusions below.

The benchmarks — where the real numbers live

This is the part that matters most, because it is where a claim becomes checkable.

Cybench (project page) is the reference for offensive, agentic capability: 40 professional CTF tasks. The published numbers are worth quoting exactly — Claude 3.5 Sonnet at 17.5% unguided, GPT-4o at 29.4% with subtask guidance — with two caveats that belong beside them every time they are cited.

They are from August 2024. That is the Cybench paper's own run, and no updated public run of current frontier models against the full suite has appeared as of August 2026. Two years is a long time in this field: treat these as the last well-documented measurement, not as today's state of the art.

And they are not comparable to each other. 17.5% is unguided; 29.4% is with subtask guidance — a materially easier setting in which the task is broken into steps for the model. Quoting them as a pair invites the reading that one model is 12 points better than the other, which is not what was measured.

What survives both caveats is the shape: whose scores these are. Frontier general models. No community security-tuned model in the survey has a published Cybench number at all — the benchmark exists, and the security models are absent from it.

CTIBench (RIT) is the mature, knowledge end: CVE-to-CWE mapping and threat-actor attribution, with the CTI-MCQA and CTI-RCM subsets that security-tuned models actually publish against — they are the numbers Cisco reports for Foundation-Sec-8B. CyberMetric (multiple-choice sets from 80 to 10,000 questions) and SecEval sit alongside it. These test what a model knows, in multiple-choice form.

CyberSOCEval targets malware analysis and threat-intel / SOC work; DefenderBench evaluates language agents in cybersecurity environments. NYU CTF Bench (200 challenges) and InterCode-CTF round out the offensive/agentic set.

The harness problem, and a harness for it

A benchmark is a set of tasks. Running it reproducibly against a model is a separate job, and it is where a lot of published numbers quietly stop being comparable — different judges, different sampling, different prompt templates, all reported as one score.

AdversariaLLM (LLM-QC/AdversariaLLM) is a direct attempt at that problem: a unified, modular toolbox for LLM robustness research implementing twelve attack algorithms against seven benchmark datasets spanning harmfulness, over-refusal and utility, with compute tracking and deterministic results as first-class features. It ships alongside JudgeZoo, which standardises thirteen judges from prior work — because "did the attack succeed" is itself a measurement, and it has been made thirteen different ways.

Its own framing of the problem is the one this post keeps arriving at: the literature is "a fragmented and oftentimes buggy ecosystem of implementations, datasets, and evaluation methods", which makes results hard to compare and progress hard to see.

A disclosure, since the name will confuse people: this toolbox and this site share a name and nothing else. It is an academic artifact from an unrelated group (arXiv 2511.04316, November 2025), it predates this site's use of the name, and it is cited here because it is the most serious answer to the measurement problem in the section about measurement.

The shape of the evidence: security knowledge benchmarks are mature and multiple-choice; offensive and agentic benchmarks are thinner and dominated by frontier general models. So the field can tell you what a model knows more readily than it can tell you what a model can do — and it can tell you almost nothing about the specific community fine-tunes most likely to be marketed to you.

What doesn't make the cut

An honest tooling survey has to name the things that are talked about but cannot be verified or downloaded:

  • CIPHER — described in the literature but its weights are withheld. A name, not a download.
  • AdvPrompter — training code (paper), no canonical checkpoint. Code, not weights.
  • PromptShield — code only.
  • RL-Hammer, PISmith, AutoInject, ArrAttack, AdvGame — research-code names with no verifiable public weights.
  • WormGPT / FraudGPT-style names — unverifiable. Forum and marketing names with no verifiable public weights behind them; treat any capability claim as unconfirmed.

None of these belong in a measured catalog as models. Some belong in it as methods — but only labeled as code, not conflated with a released weight.

The practitioner takeaway

A framework existing is not evidence that a model is good. Orchestration and attack tooling is model-agnostic by design — Garak, PyRIT, and PentestGPT run against whatever weights you supply, and none of them confer capability on the model underneath. The measurement is the benchmark, and the benchmark is exactly where the community security models are missing.

That is the gap the whole space turns on: the distance between "has a framework" and "has reproducible numbers." A project can ship a polished agent, an attack repo, and a suite of red-team probes and still have zero published, independent evaluation of the actual model it recommends.

So here is how you would actually measure a security model, and what to demand before you trust one:

  • Knowledge — CTIBench, CyberMetric, SecEval. Mature, multiple-choice, and what most security-tuned models publish against.
  • Offensive capability — Cybench, NYU CTF Bench, InterCode-CTF. Thin, and so far a general-model leaderboard.
  • Defensive / agentic — DefenderBench, CyberSOCEval.
  • Resilience and safety — Garak, PyRIT, CyberSecEval as instruments against the model, not scores for it.

If a model's card cites none of these against the specific fine-tune — not the base model, not a willingness rate, not a self-built benchmark with no third-party run — then you have a name, not a number. AdversariaLLM is being built as the catalog that keeps that line bright: the tools above are how a security model gets measured, and the measurement is the reason to catalog it at all.

For the models these instruments would be pointed at, see the 2026 survey of open security-research LLMs.