RavenX-CyberAgent 35B: A Four-Hop Repackaging With No Numbers of Its Own
D. Rose · 16 August 2026 · 6 min
RavenX-CyberAgent is a 4-bit GGUF sitting at the end of a Qwen3.6 → Opus-distilled → abliterated → security-tuned chain, marketed hard for offensive work. Every capability figure attached to it is either the base model's or self-reported on a different variant. Here is what that means before you run it.
RavenX-CyberAgent 35B is a GGUF-packaged, security-tuned build of a Qwen3.6 Mixture-of-Experts model, positioned by its author as "the most comprehensive open-source security agent model" for pentesting and bug bounty. The positioning is loud. The evidence is thin. There is no independent capability benchmark for this model, and the strong numbers circulating around it belong to a base model three hops upstream. This post separates the two.
What it actually is
The underlying architecture is Qwen3.6-35B-A3B: a Mixture-of-Experts model with 35B total parameters and roughly 3B active, 256 experts (8 routed + 1 shared), 40 layers, a 262,144-token native context, released under Apache-2.0. Because only ~3B parameters are active per token, throughput is closer to a small model than the "35B" name implies.
The RavenX repo itself ships GGUF quantizations built with llama.cpp on Apple Silicon: an F16 file at 67.8 GB and a Q4_K_M file at 20.7 GB (~4.89 bpw, marked "recommended"). So the thing you'd actually download is a 4-bit quant of a fine-tune of an abliterated model of a distilled model of Qwen.
That chain matters, so here it is in full:
Qwen3.6-35B-A3B (root MoE) → lordx64's Opus-4.7 reasoning distillation — an attention-only LoRA (r=16, ~3.44M trainable params) fine-tuned on ~7,800 reasoning traces generated via the Anthropic Claude Opus 4.7 API, teaching explicit <think>…</think> reasoning → huihui-ai's abliteration — a crude proof-of-concept refusal-removal of the distilled model → RavenX security tuning + GGUF quant — this repo.
Two clarifications the naming actively obscures. First, the "Claude-4.7-Opus" string in the repo name is inherited from upstream; the abliterated intermediate does not itself claim any Claude distillation — that claim originates one layer up at lordx64. Second, per the README metadata, the declared base_model is the huihui abliterated build — nothing else.
The maker is Gabe Garcia, Hugging Face user "deadbydawn101", whose profile bio affiliates the work with "RavenX AI Labs LLC." The profile lists 27 published models. The company's existence as a legal entity is not verifiable beyond that bio.
Makers' claims vs. what's verified
This is the section that matters most for this audience, because the gap is wide.
The card claims a "RavenX security LoRA" trained on 745K+ examples from 110 sources over 12 training rounds, with 51/51 merged LoRA tensors. None of that is verifiable: no training dataset, recipe, or evaluation is released. And the README's base_model field points only to the huihui abliterated model — so it is not independently confirmed that a distinct security LoRA is merged at all, versus this being largely a repackaging of huihui's abliterated weights.
The upstream distillation is the one part of the chain with a stated method: lordx64 documents a real LoRA and a ~7,800-trace dataset. But the faithfulness and quality of that distillation are unverified, and the abliteration and RavenX steps on top of it are not evaluated at all.
The functional claims — RATH 6-step methodology (Attack Surface, Exploit, Impact, Remediation, Document, Prevent), CVSS scoring, CWE IDs, MITRE ATT&CK mappings, kill-chain analysis, tool-calling/MCP, "autonomous agent" — are all author positioning from the card. They describe an output format the system prompt enforces, not a measured capability.
Benchmarks
There are no independently published benchmarks for this model. Every number that gets waved near it belongs to something else. The column that matters is "Measured on."
| Benchmark | Score | Measured on | Source |
|---|---|---|---|
| SWE-bench Verified | 73.4 | BASE — Qwen3.6-35B-A3B (general agentic coding) | Qwen |
| Terminal-Bench 2.0 | 51.5 | BASE — Qwen3.6-35B-A3B (general) | Qwen |
| MMLU-Pro | 85.2 | BASE — Qwen3.6-35B-A3B (general reasoning) | Qwen |
| GPQA | 86.0 | BASE — Qwen3.6-35B-A3B (general) | Qwen |
| AIME26 | 92.7 | BASE — Qwen3.6-35B-A3B (math) | Qwen |
| GSM8K CoT (flexible-extract) | 84.3% (self-reported) | INTERMEDIATE — lordx64 distilled model, not RavenX | lordx64 |
| "Overall" (Gemini 2.5 Flash as judge) | 80.9% (self-reported) | DIFFERENT VARIANT — RavenX-CyberAgent v6.2-Experimental, not this repo | v6.2 card |
| Any eval of this repo, security or general | — | none published | — |
The base-model scores are general-purpose, published by Qwen, and never measured on the abliterated + LoRA + 4-bit RavenX build. Stacking abliteration and 4-bit quantization on a model changes its behavior; none of that change is measured here. The single security-flavored figure in the whole family is that self-reported "80.9% Overall," judged by another LLM, on a different variant, with no dataset or protocol disclosed — not reproducible, and not this model.
For a practitioner
What you'd genuinely use it for. A local, non-refusing reasoning assistant for offensive-security ideation, structured into a fixed report format. Treat the base Qwen scores as the ceiling proxy for raw capability and assume degradation from the abliteration and 4-bit quant on top. The one independent hands-on account we found — Roger Gale, a BCIT cybersecurity educator, writing on Medium in June 2026 — confirms it installs via Ollama and, tested against a real organization, produced a structured low-detection reconnaissance plan and multi-phase attack-chain reasoning. Gale frames its significance not as novel attacks but as reducing friction "by turning a failed step into a revised question" for lower-skill operators. That is qualitative, single-source, and carries no numbers — but it is the only third-party evidence that exists.
How to run it. It ships as GGUF, so Ollama, LM Studio, llama.cpp, or vLLM. The Q4_K_M file is 20.7 GB on disk with ~24 GB peak RAM, feasible on a 24 GB Apple Silicon Mac; there's also an MLX 4-bit variant for Apple Silicon. The card cites "32K tested, 262K native" for context. Use the Qwen3.6 chat template with <think>…</think> reasoning blocks. Author-reported throughput on Apple Silicon is ~89 tokens/sec generation — a hardware number, not a capability one.
The real caveats. The model is abliterated: refusal-removal was applied upstream by huihui-ai, so it is designed not to refuse offensive-security requests. That is a dual-use property, and it runs fully local with no provider-side monitoring — outputs are unvetted, and no functional-validity, accuracy, or pass@k metric exists to tell you whether generated payloads actually work. Separately, on licensing: the card states Apache-2.0, consistent across the chain, but the upstream lordx64 card notes its training data was generated via Anthropic's Claude Opus 4.7 API and advises downstream users to confirm compliance with Anthropic's usage policies — an obligation the Apache tag does not resolve.
Honest limits
What is not known here is most of it:
- No independent benchmark of this repo — security or general. The only in-family security figure is self-reported, LLM-judged, and on a different variant.
- The training claims are unverifiable — 745K examples, 110 sources, 12 rounds, 51 tensors are all self-reported with nothing released.
- It is not confirmed that a distinct RavenX security LoRA is merged — the metadata points only to the abliterated model.
- Version sprawl — the repo name references v6.2,
-Experimental, and-mlxwhile the card lists v5.1 files; "which model was actually benchmarked" is ambiguous. - The company is unverified — "RavenX AI Labs LLC" exists only as an HF bio line.
- Multimodality is unconfirmed — the Qwen base card lists vision benchmarks, implying the base is multimodal, but that was not further corroborated for this build.
The honest read: this is a heavily-marketed 4-bit repackaging at the end of a four-hop chain, whose real capability is best approximated by the base Qwen model minus whatever the abliteration and quantization cost — an amount nobody has measured.
RavenX-CyberAgent 35B in the catalog: /catalog/ravenx-cyberagent-35b. For how it fits the wider field of open security-research models — and why so many of them ship the base model's numbers as their own — see the survey: Open Security-Research LLMs, 2026.