RavenX-CyberAgent 35B
Exploitation · planning · tool-calling
Offensive-security agent designed for exploitation, planning, tool-calling, and complex reasoning.
Source: deadbydawn101/RavenX-CyberAgent-Qwen3.6-35B-A3B-Opus-4.7-OpenMythos-Pentester-BugHunter-RATH-GGUF on HuggingFace
- +exploitation
- +planning
- +tool_calling
- +complex_reasoning
RavenX-CyberAgent, per its own card: an offensive-security agent tuned for exploitation, planning and tool-calling, using a 6-step "RATH" output format. Built on huihui-ai's abliterated Qwen3.6-35B-A3B base (a real, resolvable HF repo); the claimed "Opus-4.7" distillation and version numbering are vendor-stated and not independently verified.
- ›Exploitation and kill-chain planning
- ›Tool-calling / agentic offensive workflows
- ›Complex security reasoning
Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AD_API_KEY"],
base_url="https://adversariallm.ai/v1",
)
stream = client.chat.completions.create(
model="ravenx-cyberagent-35b",
messages=[{"role": "user", "content": "…"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")A GGUF-quantized 35B Mixture-of-Experts model (roughly 3B parameters active per token) fine-tuned as an autonomous security-assessment agent, served here as the RavenX-CyberAgent-35B-v5.1-Q4_K_M.gguf build (4.89 BPW, ~20.7 GB, ~24 GB peak memory). It is designed to run locally under Ollama, llama.cpp, LM Studio, and vLLM, and to drive penetration-testing and vulnerability-assessment workflows. The model structures its output using the vendor's 'RATH' six-step protocol (1-Attack Surface, 2-Exploit, 3-Impact, 4-Remediation, 5-Document, 6-Prevent), annotating findings with CVSS scores, CWE identifiers, and MITRE ATT&CK techniques. It emits Qwen-style <think> reasoning blocks with a configurable thinking depth (OFF/LOW/MED/HIGH) and supports tool-calling for common offensive tooling. The vendor markets it as 'the most comprehensive open-source security agent model in GGUF'; that framing, its claimed model lineage, and its reported metrics are not independently verified (see limitations).
Per the vendor card, the model was fine-tuned over 12 rounds on a claimed 745,000+ examples drawn from roughly 110 sources, focused on penetration testing, vulnerability assessment, and exploit development. Named components include tool-calling examples for nmap, sqlmap, nuclei, kubectl, and aws-cli; self-improving agent traces and research-automation examples; bug-bounty material (cited as ShuvonSec repositories); and offensive-security ('blackhat') conversations. These counts and sources are as stated by the vendor and are not independently verifiable.
- +Authorized security testing and penetration assessments
- +Vulnerability identification and CVSS scoring
- +Professional pentest report generation
- +Security research and education
- +Tool-calling command generation (nmap, sqlmap, nuclei, kubectl, aws-cli)
- +In-context learning from reference repository templates to match report formats
- +Multi-phase attack (kill) chain analysis
- −The card states the model is for research and authorized use ONLY: users must have explicit written authorization before assessing any target and must comply with applicable laws and regulations.
- −Because it is fine-tuned to produce offensive-security content (exploit chains, attack tooling, kill chains), a hosting harness must add its own safeguards; the card itself requires authorization checks before executing commands plus rate limiting and scope restrictions.
- −The harness must strip the model's <think> reasoning blocks from user-facing output and validate any <tool_call> JSON before execution.
- −Model lineage is only partly verifiable. The named base — 'huihui-ai/Huihui-Qwen3.6-35B-A3B-Claude-4.7-Opus-abliterated' — is a real, resolvable Hugging Face repo (an abliterated / refusal-removed Qwen3.6-35B-A3B), and the published GGUF's own header matches that lineage (architecture 'qwen35moe', ~35B MoE with ~3B active, 262K native context). What stays vendor-marketing is the rest: the 'Opus-4.7-OpenMythos' merge and any distillation from Anthropic 'Opus 4.7'/'Claude' are unverifiable, and whether 'Qwen3.6' is an official Qwen release is beyond public record — treat the distillation and version-numbering claims as unconfirmed. The base being abliterated (safety-refusals removed) is itself the reason a hosting harness must supply its own guardrails.
- −The claimed base is described as 'abliterated' (refusal-removed) and the fine-tune is offensive-security-oriented, implying reduced or removed safety guardrails; expect it to comply with harmful/unsafe requests and gate it accordingly.
- −All performance and training figures are vendor-reported throughput and training metrics on the author's own hardware, not independent third-party capability evaluations or security benchmarks.
- −Context was reportedly tested only to 32K tokens despite a claimed 262K native capacity, so long-context reliability is unverified.
The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.
Ollama Modelfile (from the card): FROM ./RavenX-CyberAgent-35B-v5.1-Q4_K_M.gguf SYSTEM "You are RavenX-Sec v5.1 by RavenX LLC. ALWAYS use EXACT 6 RATH step names: 1-Attack Surface, 2-Exploit, 3-Impact, 4-Remediation, 5-Document, 6-Prevent. Include CVSS scores, CWE IDs, and MITRE ATT&CK TTPs. Be concise. Never repeat." PARAMETER temperature 0.7 PARAMETER top_p 0.9 PARAMETER num_ctx 32768 Then: `ollama create ravenx-cyberagent -f Modelfile` and `ollama run ravenx-cyberagent`. A resulting assessment follows the RATH layout, e.g. '1-Attack Surface' (finding table with CWE-284/CWE-250/CWE-798/CWE-319), '2-Exploit' (Phase 1-5 kill chain), '3-Impact' (CVSS 9.8, full cluster compromise), '4-Remediation', '5-Document' (MITRE T1078.004, T1611, T1557), and '6-Prevent'.
Reported by the model’s authors, not our own testing — our scores are in the table above.
| Generation throughput | 89.3 tokens/sec | vendor-reported, Q4_K_M on M4 Max 128GB; not an independent capability benchmark |
| Prompt processing throughput | 900.6 tokens/sec | vendor-reported on the author's M4 Max hardware |
| Training validation loss (best) | 0.688 (Round 10) | vendor-reported training metric; loss ranged 0.674-0.926 across 12 rounds |
| Tested context length | 32K tokens | vendor-reported as tested (of a claimed 262K native capacity) |
No measurements published for this version yet.
Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.
A four-hop repackaging with no numbers of its own. The capability claims in the chain belong to models upstream of it, and no evaluation of this build has been published.
This grades the RECORD, not the model. A model rated E may be excellent — the claim is only that nobody has shown it.
RavenX-CyberAgent is a 4-bit GGUF sitting at the end of a Qwen3.6 → Opus-distilled → abliterated → security-tuned chain, marketed hard for offensive work. Every capability figure attached to it is either the base model's or self-reported on a different variant. Here is what that means before you run it.
Read the analysis →Start with RavenX-CyberAgent 35B
A confirmed email account includes 30 free messages a month.