RedSage-Qwen3-8B-DPO
Security helper model.
Source: mradermacher/RedSage-Qwen3-8B-DPO-i1-GGUF on HuggingFace
RISys-Lab's RedSage: a Qwen3-8B base preference-tuned with DPO on the Tulu-3 preference mixture, oriented toward red-team and cybersecurity assistance.
- ›Red-team planning and adversary emulation writeups
- ›Explaining offensive techniques for a defender's benefit
- ›Assisting with security tooling and scripts
Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AD_API_KEY"],
base_url="https://adversariallm.ai/v1",
)
stream = client.chat.completions.create(
model="redsage-8b",
messages=[{"role": "user", "content": "…"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")RedSage-Qwen3-8B-DPO is the final stage of RISysLab's four-stage "RedSage" cybersecurity training pipeline, built on Qwen3-8B (Qwen/Qwen3-8B-Base). The pipeline runs Continual Pre-Training, then Targeted Pre-Training, then Supervised Fine-Tuning (yielding RedSage-Qwen3-8B-Ins), and finally Direct Preference Optimization (DPO) to produce this model. The card states the DPO stage enhances general reasoning and safety while preserving the model's cybersecurity expertise. It is positioned as a general-purpose cybersecurity assistant for tasks such as log analysis, threat-intelligence summarization, and educational security queries. The card notes the model benefits from improved instruction adherence through human-preference alignment, but still cautions that outputs can be incorrect and must be verified in critical environments.
Base model: Qwen/Qwen3-8B-Base. The card describes a four-stage lineage: (1) Continual Pre-Training (CPT) -> RedSage-Qwen3-8B-CFW; (2) Targeted Pre-Training -> RedSage-Qwen3-8B-Base; (3) Supervised Fine-Tuning (SFT) -> RedSage-Qwen3-8B-Ins; (4) Direct Preference Optimization (DPO) -> this model. The DPO alignment stage was run on the AllenAI Tulu 3 Preference Mixture (dataset id allenai/llama-3.1-tulu-3-8b-preference-mixture). The card names only this DPO alignment dataset; it does not enumerate the specific cybersecurity corpora used in the earlier CPT / targeted pre-training / SFT stages.
- +General-purpose cybersecurity assistance
- +Log analysis (e.g. inspecting log entries for indicators of compromise)
- +Threat-intelligence summarization
- +Educational cybersecurity queries
- +Benefits from improved instruction adherence via human-preference (DPO) alignment
- −The model may still produce incorrect information; the card advises always verifying outputs in critical security environments
- −Despite preference alignment, outputs may contain inaccuracies and should not be relied on unverified
- −The card frames results against a 'baseline' but does not fully specify that baseline in the fetched text, so improvement deltas are the authors' own framing
The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "RISys-Lab/RedSage-Qwen3-8B-DPO"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
messages = [
{"role": "system", "content": "You are RedSage, a helpful cybersecurity assistant."},
{"role": "user", "content": "Analyze the following log entry for potential indicators of compromise: 'POST /cgi-bin/test-cgi?* HTTP/1.1'"}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))Reported by the model’s authors, not our own testing — our scores are in the table above.
| RedSage-Bench Macro Average (0-shot) | 84.83 (baseline 81.85) | vendor-reported, not independently verified; RedSage-Bench is the authors' own benchmark |
| RedSage-Bench Knowledge (General) | 82.48 (baseline 80.46) | vendor-reported, not independently verified; authors' own benchmark |
| RedSage-Bench Knowledge (Frameworks) | 83.80 (baseline 78.82) | vendor-reported, not independently verified; authors' own benchmark |
| RedSage-Bench Skill (Offensive) | 88.54 (baseline 86.16) | vendor-reported, not independently verified; authors' own benchmark |
| RedSage-Bench Tools (CLI) | 86.30 (baseline 83.92) | vendor-reported, not independently verified; authors' own benchmark |
| RedSage-Bench Tools (Kali) | 79.30 (baseline 75.56) | vendor-reported, not independently verified; authors' own benchmark |
| External Cybersecurity Benchmarks Mean (0-shot) | 81.10 (baseline 75.71) | vendor-reported, not independently verified |
| CTI-Bench (MCQ) | 70.84 (baseline 62.76) | third-party benchmark, but score is vendor-reported / not independently verified |
| CTI-Bench (RCM) | 70.60 (baseline 54.00) | third-party benchmark, but score is vendor-reported / not independently verified |
| CyberMetric | 90.00 (baseline 88.60) | third-party benchmark, but score is vendor-reported / not independently verified |
| MMLU (Security) | 79.00 (baseline 76.00) | third-party benchmark, but score is vendor-reported / not independently verified |
| SecBench (En) | 80.06 (baseline 73.26) | third-party benchmark, but score is vendor-reported / not independently verified |
| OpenLLM Leaderboard Mean | 74.33 (baseline 65.92) | vendor-reported self-run, not independently verified |
| OpenLLM MMLU | 77.07 (baseline 73.59) | vendor-reported self-run, not independently verified |
| OpenLLM ARC-C | 71.76 (baseline 62.54) | vendor-reported self-run, not independently verified |
| OpenLLM GSM8K | 82.71 (baseline 75.66) | vendor-reported self-run, not independently verified |
| OpenLLM HellaSwag | 79.87 (baseline 56.70) | vendor-reported self-run, not independently verified |
No measurements published for this version yet.
Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.
The card reports numbers on named third-party benchmarks, but every score is vendor-run and self-reported, with no method published and no independent reproduction.
This grades the RECORD, not the model. A model rated E may be excellent — the claim is only that nobody has shown it.
An imatrix GGUF quant of RISys-Lab's DPO-aligned RedSage-Qwen3-8B — an 8B cybersecurity assistant built on Qwen3-8B-Base, with an ICLR 2026 paper and a documented four-stage pipeline. The provenance is unusually clean for a community GGUF. The catch: every benchmark is author-reported, the flagship one is a benchmark the authors built themselves, no third party has reproduced any of it, and the derivative-weight license is still unresolved.
Read the analysis →Start with RedSage-Qwen3-8B-DPO
A confirmed email account includes 30 free messages a month.