# RedSage-Qwen3 8B: a paper-backed open cyber-assistant, benchmarked only by its authors

*D. Rose · 18 August 2026 · 7 min*

> An imatrix GGUF quant of RISys-Lab's DPO-aligned RedSage-Qwen3-8B — an 8B cybersecurity assistant built on Qwen3-8B-Base, with an ICLR 2026 paper and a documented four-stage pipeline. The provenance is unusually clean for a community GGUF. The catch: every benchmark is author-reported, the flagship one is a benchmark the authors built themselves, no third party has reproduced any of it, and the derivative-weight license is still unresolved.

Most of the security fine-tunes in this catalog arrive as an anonymous GGUF with a marketing README and no paper. RedSage is the other kind: an academic model with named authors, a documented training pipeline, and a paper — [arXiv 2601.22159](https://arxiv.org/abs/2601.22159), published at ICLR 2026. That earns it a closer read than a typical community upload. It also raises the bar for honesty, because a paper full of numbers is exactly the kind of thing people quote without checking who produced the numbers. Here, the answer is: the authors did, for every one of them.

The entry in this catalog is one specific build — the ["weighted/imatrix" GGUF quants](https://huggingface.co/mradermacher/RedSage-Qwen3-8B-DPO-i1-GGUF) of `RISys-Lab/RedSage-Qwen3-8B-DPO`, quantized by the prolific community packager mradermacher. That's the version you'd pull to run locally.

## What it actually is

The underlying model is an 8-billion-parameter dense, decoder-only LLM. Its foundation is [Qwen/Qwen3-8B-Base](https://huggingface.co/Qwen/Qwen3-8B-Base): 8.2B total parameters (6.95B non-embedding), 36 layers, grouped-query attention with 32 query and 8 key/value heads, a 32,768-token native context, Apache-2.0, pretrained on roughly 36 trillion tokens across 119 languages.

RedSage adapts that base through a stated four-stage pipeline, ending at the DPO checkpoint this catalog entry wraps. Per the [DPO model card](https://huggingface.co/RISys-Lab/RedSage-Qwen3-8B-DPO/raw/main/README.md), the DPO stage is trained from the instruct checkpoint using Direct Preference Optimization on `allenai/llama-3.1-tulu-3-8b-preference-mixture`. The chain, end to end:

**Qwen3-8B-Base → RedSage-CFW** (continual pre-training) **→ RedSage-Base** (targeted pre-training) **→ RedSage-Ins** (SFT) **→ RedSage-DPO** (DPO) **→ i1 GGUF** (mradermacher quantization)

The domain adaptation itself is the substance here. Per the [paper](https://arxiv.org/abs/2601.22159), the family was built from 11.8B tokens of cybersecurity continual-pretraining data (28.6K documents) plus 266K agentically-generated multi-turn SFT samples. Original weights are BF16; the GGUF you'd download is a lossy conversion of them.

The model comes from [RISys-Lab](https://huggingface.co/RISys-Lab) — the "Robust Intelligent Systems Lab." The paper's authors are Naufal Suryanto, Muzammal Naseer, Pengfei Li, Syed Talal Wasim, Jinhui Yi, Juergen Gall, Paolo Ceravolo, and Ernesto Damiani. The lead author's [page](https://naufalso.github.io/) indicates a Khalifa University affiliation, though that isn't confirmed verbatim from the paper.

## Makers' claims vs. what's verified

The authors position RedSage as "an open-source, locally deployable cybersecurity assistant with domain-aware pretraining and post-training" ([paper](https://arxiv.org/abs/2601.22159)). The DPO card frames intended use as general-purpose cybersecurity assistance, log analysis, threat-intelligence summarization, and educational queries — and, to its credit, warns that the model "may still produce incorrect information" and to "always verify outputs in critical security environments" ([model card](https://huggingface.co/RISys-Lab/RedSage-Qwen3-8B-DPO)). The stated motivation is privacy: running security workflows without shipping sensitive data to a proprietary API.

So the provenance is clean and the intent is documented. The honest split is about *who measured the results*: every RedSage benchmark below is author-reported, and the flagship number comes from a benchmark the authors built themselves. No independent, third-party evaluation of any RedSage variant was found. That doesn't make the numbers wrong — it makes them unreplicated. Read them as "the authors' own measurement of the authors' own model," which is the correct frame for a fresh research release that no one else has re-run yet.

One thing that separates RedSage from the abliterated offensive-security builds elsewhere in this catalog: it is not refusal-stripped. It's DPO-aligned, and the card itself tells you to check its output. That's a defensive-assistant posture, not a jailbroken one.

## Benchmarks

Every number here is **author-reported**, for the **full-precision (BF16) RedSage-Qwen3-8B-DPO** — not the GGUF quant in this catalog, and not independently reproduced. The "author's baseline" column is the comparator reported alongside each score; the paper's results tables didn't render cleanly enough to confirm which model the baseline is, but it is most likely stock Qwen3-8B (unconfirmed).

| Benchmark | RedSage-DPO | Author's baseline | Which model / caveat |
|---|---|---|---|
| RedSage-Bench (0-shot, macro avg) — *the authors' own benchmark* | 84.83 | 81.85 | RedSage-DPO, author-reported |
| External cybersecurity mean (CTI-Bench, CyberMetric, MMLU-Security / SECURE) | 81.10 | 75.71 | RedSage-DPO, author-reported |
| Open LLM Leaderboard (mean) | 74.33 | 65.92 | RedSage-DPO, author-reported |

Source: the [DPO model card](https://huggingface.co/RISys-Lab/RedSage-Qwen3-8B-DPO). The paper's abstract states the headline a different way — RedSage 8B "surpasses baselines by up to +5.59 points on cybersecurity benchmarks and +5.05 points on Open LLM Leaderboard tasks" ([paper](https://arxiv.org/abs/2601.22159)) — a family-level claim, so don't pin it to the DPO stage specifically.

Two caveats on the table itself. First, the flagship row is the authors' *own* benchmark (RedSage-Bench); a model scoring well on a test its makers designed is the least surprising result in the table, and the more portable signal is the external-benchmark row against public suites like CTI-Bench and CyberMetric. Second, the exact decimals were retrieved via a page-summarizing fetch rather than read cell-by-cell from a rendered table — treat them as high-confidence but verify against the card before you quote them anywhere load-bearing.

There is no clean base-model row to compare against. Qwen3-8B-Base's own published benchmark scores weren't captured for this write-up — the Qwen card defers them to a separate report — so the only base-vs-fine-tune signal available is the authors' internal "baseline" comparator above, not an independent measurement of the foundation model.

## For a practitioner

**What you'd genuinely use it for.** A local cybersecurity assistant for the drudge tasks the authors target: log triage and explanation, threat-intelligence summarization, and educational/explanatory queries — on hardware you control, with nothing leaving the box. Treat that as the documented *intent*, not a proven capability, and keep the card's own warning in the loop: verify anything that feeds a real decision.

**How to run it.** These are GGUF weights for llama.cpp, Ollama, and LM Studio. This repo is the imatrix ("i1") line — quants weighted with an importance matrix, generally higher quality than plain static quants at the same bitrate, especially at low bit-widths. The card's practical picks are `i1-Q4_K_M` (~5.03 GB, flagged "fast, recommended") or `i1-Q4_K_S` (~4.8 GB, "optimal size/speed/quality"); the range runs from `IQ1_S` (~2.12 GB, expect real quality loss) up to `Q6_K` (~6.73 GB) ([quant repo](https://huggingface.co/mradermacher/RedSage-Qwen3-8B-DPO-i1-GGUF)). At that footprint it fits comfortably on an 8 GB GPU, or on CPU. If you'd rather have non-imatrix builds, the static-quant sibling repo is `mradermacher/RedSage-Qwen3-8B-DPO-GGUF`.

**Prompt format.** The instruct/DPO lineage uses ChatML (`<|im_start|>` / `<|im_end|>`), inherited from Qwen3's instruct format and confirmed on the [instruct card](https://huggingface.co/RISys-Lab/RedSage-Qwen3-8B-Ins). Make sure your runtime applies the ChatML template rather than a generic one.

**Context length.** Neither the DPO card nor the GGUF card restates it. The Qwen3-8B-Base foundation is 32,768 tokens natively ([Qwen3-8B-Base](https://huggingface.co/Qwen/Qwen3-8B-Base)), which the quant inherits unless your runtime caps it lower — so confirm the setting rather than assuming the full 32K.

**The caveat that will actually bite you: you're benchmarking a model you're not running.** Those scores are for the BF16 original. The file you download is a [lossy quant](/notes/what-quantization-costs), and at the low-bpw end (`IQ1_S` through the Q4 picks) the degradation is real and unmeasured — no quant-specific numbers are published for any of these builds. If accuracy matters more than footprint, stay at the higher quants.

**Licensing — the genuinely unresolved part.** The base Qwen3-8B-Base is cleanly Apache-2.0 ([Qwen3-8B-Base](https://huggingface.co/Qwen/Qwen3-8B-Base)). The *derivative* RedSage weights are not licensed on the card at all — the DPO HF page shows no populated license field, and mradermacher's GGUF repo adds no grant of its own. A [GitHub issue (#15)](https://github.com/RISys-Lab/RedSage/issues/15), opened 2026-07-16, asks RISys-Lab to clarify the weight license; as of writing it's still open with no maintainer response. Until they answer, treat redistribution and commercial terms as unresolved, even though the foundation model is permissive.

## Honest limits

- **No independent evaluation exists.** Every RedSage number is author-reported, and the flagship RedSage-Bench is the authors' own benchmark. Nobody outside the lab has reproduced any of it yet.
- **The baseline comparator is unconfirmed.** The paper's results tables didn't render through fetch, so which model "baseline" refers to (most likely stock Qwen3-8B) is inferred, not verified.
- **No base-model scores here.** Qwen3-8B-Base's own benchmarks weren't captured, so there's no clean independent base-vs-fine-tune comparison — only the authors' internal delta.
- **Exact decimals are fetch-derived.** The card figures came through a page summarizer, not a cell-by-cell read; verify before quoting.
- **Context length and author affiliation are inferred.** The 32K window is the base model's native context assumed to carry through the quant; the Khalifa University affiliation comes from the lead author's page, not the paper text.
- **The derivative-weight license is a real gap**, not just a research gap — an open, unanswered licensing question, not a formality.

---

RedSage is one of the more credible-looking open security models to appear recently: a real paper, a documented four-stage pipeline, named academic authors, and an 8B footprint that runs on a laptop. The discipline it demands is the same as everything else in a measured catalog — keep the labels straight. Author-reported, not independent. Own benchmark, not a public one. BF16 in the paper, 4-bit on your disk. And a license that, for the weights that matter, nobody has actually stated yet.

See the catalog entry: [/catalog/redsage-8b](/catalog/redsage-8b). For where it sits among the rest of the field, see the [2026 survey of open security-research LLMs](/blog/open-security-research-llms-2026).
