CyberStrike-OffSec-35B

pentest planning · tool-calling · agent routing

A LoRA-aligned offensive-security agent on Qwen3.6-35B-A3B (a Mixture-of-Experts base) that emits structured Qwen XML tool calls for a pentesting harness — it plans and routes, it does not execute.

Context8,192 (unconfirmed)
Curationcurated
AvailabilityComing soon
Price / Mtoknot yet priced
BaseQwen3.6-35B-A3B
QuantizationQ4_K_M
License—

Source: oyildirim/CyberStrike-OffSec-35B-GGUF on HuggingFace

Good for
  • +exploitation
  • +pentest_planning
  • +tool_calling
  • +agent_routing
Model cardfrom the published weights — parameters, architecture, training, intended use
Parameters35B MoE (Qwen3 A3B — ~3B active per token)
ArchitectureQwen3 MoE (Qwen3.6-35B-A3B base) — LoRA-aligned
Modalitytext
Training

CyberStrike-OffSec, per its card: a single-epoch LoRA alignment on a 300-example multi-turn tool-call SFT set (held-out token accuracy 98.5%, author-reported) over the Qwen3.6-35B-A3B base, to emit structured Qwen XML tool calls for an offensive-security harness. The author frames it as a small behavioural alignment, not new capability.

Intended use
  • ›Authorized penetration-testing planning
  • ›Structured tool-call generation (Qwen XML <tool_call>) for a harness
  • ›Agent/archetype routing (web-application, explore, task)
Full model cardcurated from the model's HuggingFace card — description, training, usage, limitations, reported benchmarks

A LoRA-aligned offensive-security agent built on Qwen3.6-35B-A3B — a Qwen3 Mixture-of-Experts base (~35B total, ~3B active per token) with GDN linear-attention. It is tuned to emit structured Qwen XML tool calls for a penetration-testing harness (archetypes such as web-application, explore, and task) and to route between them, terminating cleanly; it reasons about methodology and proposes actions but does not itself execute anything. The author is explicit that this is a small, targeted behavioural alignment, not a capability upgrade over the base.

Training data

Per the model card: a single-epoch LoRA fine-tune (r=32, α=64; 42.3M trainable parameters across 310 modules covering full-attention q/k/v/o, the GDN linear-attention path, and MLP) over a 300-example set of multi-turn tool-call SFT examples, with a held-out token accuracy of 98.5%. The corpus source is not specified on the card, and the figures are author-reported, not independently verified.

Intended use
  • +Authorized offensive-security testing and research only
  • +Structured tool-call generation for a pentesting harness (Qwen XML <tool_call>)
  • +Agent/archetype routing (web-application, explore, task)
  • +Penetration-testing methodology and plan drafting
Limitations
  • −The card states the model is for authorized offensive-security testing and research ONLY: users are responsible for operating solely against systems they are authorized to test.
  • −It plans and emits tool calls; it does not execute. A hosting harness must validate any <tool_call> before acting on it and enforce its own scope and authorization checks.
  • −The fine-tune is a thin LoRA alignment (300 examples, one epoch) that the author explicitly calls a behavioural alignment, not a capability upgrade — broad capability comes from the Qwen3.6-35B-A3B base, whose lineage is beyond our verification.
  • −License is listed only as "other" with no explicit text on the card; read the source repo's terms before any use.
  • −Not served on this platform. The canonical checkpoint ships safetensors only (no first-party GGUF), and its base architecture (qwen35moe) is the same one our current worker image cannot load — so serving waits on worker support for the architecture (the same blocker as RavenX). An earlier -GGUF build is deprecated and gated by its author ("do not use for new deployments").
  • −Reported scores are the author's own 24-scenario evaluation, not an independent third-party benchmark.
Running the weights yourself

The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.

Chat use is agentic: you provide a system prompt that authorizes and scopes the engagement, plus the tool schemas, and the model replies with structured tool calls for your harness to validate and run.

System (example): "You are an offensive-security planning assistant. Only propose actions against assets the user states they are authorized to test. Emit each action as a Qwen XML <tool_call> and stop; never claim a result you have not been given as an observation."

User: "Authorized web-app assessment of https://staging.example.com — start the recon phase."

The model routes to the web-application archetype and emits <tool_call> XML (e.g. an nmap or a directory-enumeration call) for the harness to validate and execute, rather than describing the commands in prose.
Benchmarks — as reported on the card

Reported by the model’s authors, not our own testing — our scores are in the table above.

Genuine structured tool calls18/24author's 24-scenario eval (6 axes × difficulty; ~62% out-of-distribution); not independent
Correct tool/archetype routing10/24author-reported on the same 24-scenario eval
Clean termination24/24author-reported
Fabricated observations0/24author-reported — the model did not invent tool results
Scores

No measurements published for this version yet.

Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.

Versionsscores attach to a version; v2 does not inherit v1's numbers
v12026-08-11—current

Start with CyberStrike-OffSec-35B

A confirmed email account includes 30 free messages a month.

Sign in to start
CyberStrike-OffSec-35B — Offensive security agent · AdversariaLLM