Open security-research LLMs in 2026: a measured field survey

D. Rose · 31 July 2026 · Updated 31 August 2026 · 12 min

The open security-LLM field has two camps and a rigor gap between them — institutions that publish reports and benchmarks, and community "uncensored" builds that publish almost nothing you can check. A cited tour of both, the shared base models underneath them, the benchmarks that actually exist, and why the offensive leaderboards are still led by general models.

There are more open cybersecurity language models than there were a year ago, and fewer of them are measured than the marketing implies. The field is real and growing. It is also under-measured, and the gap between the two is the whole story.

It splits into two camps. On one side are institution- and academic-backed models that publish technical reports, name their base models, and report scores on shared benchmarks — Cisco's Foundation-Sec, Trend Micro's Primus, RISys-Lab's RedSage, IBM's CyberPal, Cisco's Antares. On the other are community "uncensored" offensive builds positioned by marketing language — "uncensored red team," "745K+ examples," "180 languages" — that publish essentially no independent evaluation: WhiteRabbitNeo/DeepHat, BugTraceAI, HIDra, HEX, Dolphin3-Cyber, and dozens of GGUF fine-tunes like them. Almost every model in both camps is a continued-pretraining run or a fine-tune of the same handful of open base models. The security tuning is real. The measurement mostly is not.

The shared foundations

Start with what is underneath, because it explains most of the confusion. Nearly all of these "security models" are thin tunes of four open base-model families:

  • Llama 3.1 (8B / 70B, and the Nemotron 70B variant) — base for Foundation-Sec, Llama-Primus, Dolphin3-Cyber, and the claimed base for WhiteRabbitNeo V2.
  • Qwen (2.5, 3, 3.6; dense and the newer A3B mixture-of-experts) — base for RedSage, HIDra, HEX-30B, CyberStrike-OffSec, RavenX, several BugTraceAI variants, and a wave of "uncensored" Qwen-Coder builds.
  • DeepSeek-Coder (6.7B / 33B) — base for the WhiteRabbitNeo code line.
  • Gemma (now Gemma 4, MoE) — base for BugTraceAI-Apex.

The consequence is direct: when a security model is a light fine-tune of Llama 3.1 or Qwen, the only real benchmark number often available is the base model's — and quoting it as the fine-tune's capability is exactly the fabrication a measured catalog has to refuse. Throughout this survey, and in every catalog deep-dive it links, base-model numbers are labeled as the base's and never as the security build's.

Camp one: the institutions that publish

These models name a base, ship a report, and report on shared benchmarks. Their reported gains are modest and narrow — mostly single-digit to roughly 15% aggregate improvement, and mostly on knowledge and CTI multiple-choice benchmarks — but they are citable, which is the point.

ModelCreatorBaseWhat is published
Foundation-Sec-8BCisco Foundation AILlama 3.1 8B (continued pretraining)Base, Instruct, and Reasoning reports; vendor CTI benchmarks
Primus / Llama-PrimusTrend MicroLlama 3.1 8B / Nemotron-70BTechnical report + open datasets; aggregate-gain numbers
RedSage-Qwen3-8BRISys-LabQwen3-8B-BaseICLR 2026 paper; author-reported benchmarks (incl. their own)
CyberPal2.0IBM Research + collaboratorsgpt-oss 20B-derivedPaper + SecKnowledge-Eval dataset
AntaresCisco Foundation AIGranite 4.0 (350M / 1B)Technical report + a vulnerability-localization benchmark
ZySec-7BZySec-AIZephyr / Mistral 7BApache-2.0 weights; no independent benchmark

Foundation-Sec is the best-evidenced entry in the whole field. Cisco built it on Llama-3.1-8B via continued pretraining, then released a base (arXiv 2504.21039), an instruct variant, and a reasoning model trained with SFT plus reinforcement learning from verifiable rewards (arXiv 2601.21051). It has real CTI numbers — CTI-MCQA 0.691 vs Llama-3.1-8B's 0.607, CTI-RCM 0.753 vs 0.531, CTI-VSP 0.856 vs 0.811 (model card). Two hedges keep it honest: those scores are Cisco's own with no independent third-party reproduction, and the catalog build is the lossy Q4_K_M 4-bit GGUF of the reasoning model, which nobody has separately measured against the full-precision numbers. Licensing on the reasoning derivative is not a clean single tag — read its NOTICE.md before relying on any one label.

Primus / Cybertron (Trend Micro) is the open-dataset counterpart. Llama-Primus-Base is continued pretraining of Llama-3.1-8B-Instruct on 2.77B security tokens, reported at +15.88% aggregate over benchmarks; the Nemotron-70B variant adds 7.6B tokens for a reported +11.19% (arXiv 2502.11191, EMNLP 2025). Evaluated on CTI-Bench, CyberMetric, and SecEval, with the Primus-Seed / FineWeb / Instruct / Reasoning datasets released. These are Trend Micro's own reported numbers — real, but narrow, and on knowledge/CTI multiple-choice rather than hands-on tasks.

RedSage (RISys-Lab) has the cleanest provenance of any community-distributed GGUF here: a documented four-stage pipeline on Qwen3-8B-Base, published at ICLR 2026 (arXiv 2601.22159). It reports 81.10 on external cyber benchmarks (vs 75.71 baseline) and 74.33 on the Open LLM Leaderboard (vs 65.92). The catch is structural: every number is author-reported, the flagship RedSage-Bench (84.83) is a benchmark the authors built themselves, no third party has reproduced any of it, and the derivative-weight license is still unresolved on an open GitHub issue.

CyberPal2.0 (IBM Research and collaborators) is the public gpt-oss-20B-derived defensive model — CTI, SOC, incident response, CVE-to-CWE mapping — with a paper (arXiv 2510.14113) and the SecKnowledge-Eval dataset; the paper also evaluates Qwen3 4B/8B/14B variants that were not public when checked. Antares (Cisco Foundation AI) is a smaller, more specialized bet: Granite-4.0-based 350M and 1B models for agentic vulnerability localization from a terminal, with a technical report and its own vulnerability-localization benchmark (a described 3B is not released). ZySec-7B is the community-permissive edge of this camp — a Zephyr/Mistral-based, Apache-2.0 defensive assistant spanning 30+ security domains (model card) — but, like most of the camp, it publishes no independent third-party benchmark at all.

A few more sit in this "not advertised as unconditionally uncensored" group with lighter documentation: CyberBase-13B (CyberNative, a Llama-2/Vicuna-13B experimental continued-training base, model card only), CyberStrike-OffSec (a Qwen3.6-35B-A3B MoE tool-calling pentest build, no paper), and BaronLLM (Alican Kiraz, a gated Llama-3.1-8B offensive-simulation model that is explicitly safety-aligned rather than uncensored — its card carries an inconsistent 70B example against an 8B repo, so treat details as low-confidence).

Camp two: the community "uncensored" builds

This camp sells the removal of refusals as the feature. "Uncensored" is the publisher's description, not an independently verified property, and the pattern is consistent: an abliterated or refusal-suppressed fine-tune of an open base, distributed as GGUF, documented by a model card, with no paper and no reproducible capability number.

ModelBase (publisher-stated)Published eval of the fine-tune?
WhiteRabbitNeo / DeepHatDeepSeek-Coder, CodeLlama, Llama 3/3.1, Qwen2.5-CoderNone; card tunes the model not to refuse cyber requests
BugTraceAI-Apex-G4Gemma 4 26B A4B MoE (via abliteration)Self-reported compliance rate only
RavenX-CyberAgent-35BQwen3.6-35B-A3B (distill + abliteration)None for this repo
HIDra-30Babliterated Qwen3-Coder-30B-A3BNone (card only)
HEX-30Babliterated Qwen3-30B-A3BNone (card only)
Dolphin3-CyberDolphin3.0 / Llama-3.1-8B (abliterated)None (card only)

WhiteRabbitNeo (the Kindo-backed project, now also branded DeepHat) is the anchor of this camp — an offensive+defensive series across DeepSeek-Coder 6.7B/33B, CodeLlama 13B, Llama 3/3.1 8B/70B, and Qwen2.5-Coder 7B, whose cards instruct the model not to refuse cyber requests. The catalog holds a community GGUF of WhiteRabbitNeo-33B-v1.5, built on DeepSeek-Coder-33B (model card). The security tuning is real; the evidence for it is not — no benchmark of the fine-tune exists, so the only citable numbers are DeepSeek-Coder's base coding scores. One trap: Kindo's "~5% of Cybench tasks solved" figure is for the separate V3 line, not this 33B v1.5, and must not be attributed to it. License is the DeepSeek Coder agreement plus a use-restricted WhiteRabbitNeo extension (the 13B-v1 carries a Llama-2-plus-restrictions custom license) — check it before any commercial deployment.

BugTraceAI-Apex-G4-26B is the clearest illustration of the willing-versus-capable gap. It is a DPO-tuned, deliberately uncensored build reached from Google's Gemma 4 26B A4B mixture-of-experts (25.2B total / 3.8B active) via an intermediate abliterated model, shipped by an anonymous Hugging Face account (model card). The one number attached to the fine-tune is a self-reported 100% compliance / 0% refusal on a CyberSecEval MITRE ATT&CK set — a measure of whether the model answers, not whether its exploits are correct or functional. The strong figures you will see (MMLU-Pro 82.6% and the rest) are Gemma 4's base-model general scores, not this build's security capability. The broader BugTraceAI family (CORE-Fast on Qwen2.5-Coder 7B, CORE-Pro on Mistral-Nemo 12B, CORE-Ultra on Qwen3.6 27B) is self-reported throughout, with no paper.

RavenX-CyberAgent-35B is the thinnest-evidenced catalog entry: a 4-bit GGUF at the end of a four-hop chain — Qwen3.6-35B-A3B, distilled on Claude Opus 4.7 reasoning traces by one author, abliterated by another, then security-tuned and repackaged by a solo maker. Its headline numbers (SWE-bench 73.4, MMLU-Pro 85.2, GPQA 86.0) are the Qwen3.6 base model's official general scores, never measured on the abliterated-plus-LoRA-plus-4-bit build; the only security-flavored figure in the family (80.9%, LLM-judged) is self-reported on a different variant, and the sole independent coverage is one practitioner's qualitative hands-on account.

The long tail of this camp follows the same template and is worth naming so the pattern is unmistakable: HIDra (an abliterated Qwen3-Coder-30B-A3B for red-team/exploitation/BadUSB work, card only), HEX-30B (an abliterated Qwen3-30B-A3B SFT merge, no paper), Dolphin3-Cyber (an abliterated Dolphin3.0/Llama-3.1-8B, card only), Qwen3.6-27B-Uncensored-Cyber (a refusal-suppressed dense Qwen3.6-27B whose own card states it adds no new cyber knowledge), KiKai (an Indonesian/English LoRA over DeepHat-V1-7B, no paper), Navi (an uncensored Llama-3.2-3B llamafile), Lily Cybersecurity Uncensored (an abliterated Mistral-7B GGUF, no paper), and qwen25_UNCENSORED (a merged-plus-LoRA Qwen2.5-Coder 7B, whose one linked paper is an MDPI article). None publishes an accuracy, pass@k, or functional-validity number for its offensive output.

The tension: willing versus capable

Willingness and capability are different axes, and the "uncensored" marketing conflates them. A model that answers every exploit request tells you it will not refuse. It does not tell you the exploit works. Not one of the community builds above publishes an accuracy or functional-validity number for its offensive output. The institution camp is the inverse: more honest, less exciting, reporting modest gains on knowledge benchmarks that say little about live triage, exploit development, or malware reverse-engineering. Neither camp has shown, in public, that its security tuning beats a general model at a real task.

Where the honest evaluation actually is

Citable evaluation exists; it is just unevenly distributed, and it mostly does not touch the models above.

Security knowledge is the mature end. CTIBench (RIT — CVE-to-CWE mapping, threat-actor attribution; its CTI-MCQA and CTI-RCM subsets are what Cisco reports against), CyberMetric, and SecEval are the benchmarks security-tuned models publish numbers on. They test what a model knows, in multiple-choice form.

Offensive and agentic capability is thinner — and led by general models. Cybench (40 professional CTF tasks) is the reference, and its top scorers are frontier general models: Claude 3.5 Sonnet at 17.5% unguided and GPT-4o at 29.4% with subtask guidance — figures from the Cybench paper's own run in August 2024, and not comparable to each other, since the second is measured in the easier guided setting. No updated public run against the full suite has appeared as of August 2026. DefenderBench evaluates language agents in cybersecurity environments; CyberSOCEval adds malware-analysis and SOC-oriented tasks; Meta's CyberSecEval / Purple Llama suite covers insecure-code generation, injection, and offensive-compliance risk.

Here is the gap that matters. These offensive and agentic evals almost never test the community security-tuned models — they benchmark frontier general models. There is no public head-to-head showing that WhiteRabbitNeo, ZySec, or even Foundation-Sec beats a general model at a real CTF or a live pentest. So the field has solid benchmarks for security knowledge, thin ones for offensive capability, and essentially no rigorous public evaluation of the specific fine-tunes most likely to be marketed to a practitioner. That absence is the survey's central finding, not a footnote to it. The benchmarks, the attack algorithms (GCG, PAIR, TAP, AutoDAN, Rainbow Teaming), and the orchestration frameworks (PentestGPT, Garak, PyRIT) get their own treatment in Security-LLM frameworks and benchmarks.

The defensive and tooling layer

Around the security-tuned models sits a much better-measured ecosystem the platform also tracks — and, unlike the offensive camp, most of it ships peer-reviewed papers.

Attack generators produce adversarial prompts for resilience testing rather than answering cyber questions: AmpleGCG (gated Llama-2 generators, arXiv 2404.07921), ADV-LLM (six self-tuned suffix generators, arXiv 2410.18469), Self-RedTeam (self-play attacker/defender, arXiv 2506.07468), and Jailbreak-R1 (a gated RL generator). Prompt-injection detectors are the most measured corner of all: Meta's Prompt Guard (mDeBERTa, 86M/22M), PIGuard/InjecGuard (trained against over-defense false positives, ACL 2025), and Sentinel (arXiv 2506.05446) — covered in depth in Defending against prompt injection. Broader safety guards classify harmful prompts and responses: Llama Guard (arXiv 2312.06674), Qwen3Guard (arXiv 2510.14276), IBM's Granite Guardian (arXiv 2412.07724), and Google's ShieldGemma (arXiv 2407.21772). And a layer of cyber encoders powers CTI pipelines — SecureBERT (arXiv 2204.02685), the Cisco SecureBERT 2.0 ModernBERT rebuild (arXiv 2510.00240), CySecBERT, and ATT&CK-mapping classifiers like TTPXHunter.

The platform owner's deployment recommendation follows straight from this map: for serious use, layer a dedicated injection detector (Prompt Guard 2, PIGuard, Sentinel), a separate content-policy guard (Qwen3Guard, Llama Guard, Granite Guardian), and task-specific cyber models kept behind tool permissions, sandboxing, and audit logs. No single security-tuned model is the safety layer.

The rest of the catalog: helpers, not security models

Three catalog entries are general models with no security-specific evaluation, present as helpers. Qwen2.5-14B is Alibaba's Apache-2.0 14.7B dense base checkpoint — strong self-reported base benchmarks (MMLU 79.7, GSM8K 90.2), but the card is explicit it is "not recommended for conversations" out of the box, and there is no security eval of it whatsoever. LLaVA-1.5-7B (CLIP ViT-L/14 on a Vicuna-7B backbone, peer-reviewed CVPR 2024 benchmarks) and MiniCPM-V first-gen (a ~3B bilingual VLM with self-reported card numbers) are vision helpers; any security use is inference-time repurposing, not a documented capability. Note the attribution trap on MiniCPM-V: the "GPT-4V level MLLM on your phone" paper is about a later 2.5 model, not this first-gen 3B, and its frontier-beating claims must not be carried over.

How to read any entry

The pattern across the field is consistent enough to make a checklist. Before you trust a number:

  • Separate the base from the fine-tune. If a number is not tied to a named evaluation of the specific fine-tune, treat it as the base model's. Most circulating figures are.
  • Separate willing from capable. A compliance or refusal rate measures behavior, not exploit quality. Ask for accuracy, pass@k, or functional validity — and note when it is absent, which for the uncensored camp is always.
  • Separate self-reported from independent. Vendor and author numbers are worth reading; they are not third-party verification, and almost none of these models have any.
  • Check the actual build. The distributed artifact is usually a quantized GGUF, not the full-precision model the benchmarks describe. Quantization loss is generally unmeasured.
  • Read the license, not the tag. "Open-weight" is not OSI-open. Several models here are gated, research-only, or under use-restricted custom terms; some have licenses that are unresolved on an open issue. A single "apache-2.0" line does not settle it.

That checklist is the product. A measured catalog is not one that promises these models beat a general LLM at your job — the public evidence supports that claim for none of them. It is one where every number is traced to its source, the base model and the fine-tune are never conflated, and the gaps are labeled as plainly as the scores. That is what AdversariaLLM is built to be: the open security-research field, measured. The field is real and growing. The measurement is the reason to bring it in at all.