CyberStrike-OffSec-35B
pentest planning · tool-calling · agent routing
A LoRA-aligned offensive-security agent on Qwen3.6-35B-A3B (a Mixture-of-Experts base) that emits structured Qwen XML tool calls for a pentesting harness — it plans and routes, it does not execute.
Source: oyildirim/CyberStrike-OffSec-35B-GGUF on HuggingFace
- +exploitation
- +pentest_planning
- +tool_calling
- +agent_routing
CyberStrike-OffSec, per its card: a single-epoch LoRA alignment on a 300-example multi-turn tool-call SFT set (held-out token accuracy 98.5%, author-reported) over the Qwen3.6-35B-A3B base, to emit structured Qwen XML tool calls for an offensive-security harness. The author frames it as a small behavioural alignment, not new capability.
- ›Authorized penetration-testing planning
- ›Structured tool-call generation (Qwen XML <tool_call>) for a harness
- ›Agent/archetype routing (web-application, explore, task)
A LoRA-aligned offensive-security agent built on Qwen3.6-35B-A3B — a Qwen3 Mixture-of-Experts base (~35B total, ~3B active per token) with GDN linear-attention. It is tuned to emit structured Qwen XML tool calls for a penetration-testing harness (archetypes such as web-application, explore, and task) and to route between them, terminating cleanly; it reasons about methodology and proposes actions but does not itself execute anything. The author is explicit that this is a small, targeted behavioural alignment, not a capability upgrade over the base.
Per the model card: a single-epoch LoRA fine-tune (r=32, α=64; 42.3M trainable parameters across 310 modules covering full-attention q/k/v/o, the GDN linear-attention path, and MLP) over a 300-example set of multi-turn tool-call SFT examples, with a held-out token accuracy of 98.5%. The corpus source is not specified on the card, and the figures are author-reported, not independently verified.
- +Authorized offensive-security testing and research only
- +Structured tool-call generation for a pentesting harness (Qwen XML <tool_call>)
- +Agent/archetype routing (web-application, explore, task)
- +Penetration-testing methodology and plan drafting
- −The card states the model is for authorized offensive-security testing and research ONLY: users are responsible for operating solely against systems they are authorized to test.
- −It plans and emits tool calls; it does not execute. A hosting harness must validate any <tool_call> before acting on it and enforce its own scope and authorization checks.
- −The fine-tune is a thin LoRA alignment (300 examples, one epoch) that the author explicitly calls a behavioural alignment, not a capability upgrade — broad capability comes from the Qwen3.6-35B-A3B base, whose lineage is beyond our verification.
- −License is listed only as "other" with no explicit text on the card; read the source repo's terms before any use.
- −Not served on this platform. The canonical checkpoint ships safetensors only (no first-party GGUF), and its base architecture (qwen35moe) is the same one our current worker image cannot load — so serving waits on worker support for the architecture (the same blocker as RavenX). An earlier -GGUF build is deprecated and gated by its author ("do not use for new deployments").
- −Reported scores are the author's own 24-scenario evaluation, not an independent third-party benchmark.
The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.
Chat use is agentic: you provide a system prompt that authorizes and scopes the engagement, plus the tool schemas, and the model replies with structured tool calls for your harness to validate and run. System (example): "You are an offensive-security planning assistant. Only propose actions against assets the user states they are authorized to test. Emit each action as a Qwen XML <tool_call> and stop; never claim a result you have not been given as an observation." User: "Authorized web-app assessment of https://staging.example.com — start the recon phase." The model routes to the web-application archetype and emits <tool_call> XML (e.g. an nmap or a directory-enumeration call) for the harness to validate and execute, rather than describing the commands in prose.
Reported by the model’s authors, not our own testing — our scores are in the table above.
| Genuine structured tool calls | 18/24 | author's 24-scenario eval (6 axes × difficulty; ~62% out-of-distribution); not independent |
| Correct tool/archetype routing | 10/24 | author-reported on the same 24-scenario eval |
| Clean termination | 24/24 | author-reported |
| Fabricated observations | 0/24 | author-reported — the model did not invent tool results |
No measurements published for this version yet.
Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.
Start with CyberStrike-OffSec-35B
A confirmed email account includes 30 free messages a month.