Research

How these systems work — and how they break.

In-depth, cited research: teardowns, disclosures, and analysis from putting these models against real security work. Rigorous where it counts, and honest about what each finding does and doesn’t show.

26 posts· latest 27 September 2026

All posts

26 of 26 posts
More
18 August 20267 minD. Rose

RedSage-Qwen3 8B: a paper-backed open cyber-assistant, benchmarked only by its authors

An imatrix GGUF quant of RISys-Lab's DPO-aligned RedSage-Qwen3-8B — an 8B cybersecurity assistant built on Qwen3-8B-Base, with an ICLR 2026 paper and a documented four-stage pipeline. The provenance is unusually clean for a community GGUF. The catch: every benchmark is author-reported, the flagship one is a benchmark the authors built themselves, no third party has reproduced any of it, and the derivative-weight license is still unresolved.

16 August 20266 minD. Rose

RavenX-CyberAgent 35B: A Four-Hop Repackaging With No Numbers of Its Own

RavenX-CyberAgent is a 4-bit GGUF sitting at the end of a Qwen3.6 → Opus-distilled → abliterated → security-tuned chain, marketed hard for offensive work. Every capability figure attached to it is either the base model's or self-reported on a different variant. Here is what that means before you run it.

16 August 20267 minD. Rose

Defending against prompt injection: detectors, guards, and the difference

The open-weight defensive layer is real and growing, but it is three different jobs that get sold as one. A cited tour of the dedicated injection detectors, the broader content-policy guards, and the evaluation judges that are not runtime filters — plus the attack generators they are tested against, and how the pieces actually stack in a deployment.

15 August 202629 minD. Rose

The COLDCARD RNG Vulnerability: How a “Random” Bitcoin Seed Became Guessable

A plain-language teardown of the 2026 COLDCARD RNG flaw: how a firmware integration mistake pushed wallet generation onto a deterministic software PRNG, why the public blockchain let attackers verify guesses for free, why an air-gapped device did not help, and why a firmware update cannot repair a seed that was already generated.

12 August 20267 minD. Rose

WhiteRabbitNeo 33B: A DeepSeek-Coder Security Fine-Tune With No Benchmarks of Its Own

A cybersecurity fine-tune of DeepSeek-Coder-33B, packaged here as a community GGUF and tuned not to refuse offensive-security questions. The security tuning is real; the independent evidence for it is not — no one has published a single benchmark of the fine-tune itself. Here is what that means and how to run it.

31 July 20268 minD. Rose

The tooling around security LLMs: what orchestrates, what attacks, and what actually measures

The frameworks, attack algorithms, and benchmarks in this space are not security models — they run against whatever weights you point them at. A cited tour of the pentest agents, red-team scanners, jailbreak algorithms, and the evaluations that produce real numbers (Cybench, CTIBench, CyberSOCEval, DefenderBench) — and why "has a framework" is not "has reproducible numbers."

31 July 202612 minD. Rose

Open security-research LLMs in 2026: a measured field survey

The open security-LLM field has two camps and a rigor gap between them — institutions that publish reports and benchmarks, and community "uncensored" builds that publish almost nothing you can check. A cited tour of both, the shared base models underneath them, the benchmarks that actually exist, and why the offensive leaderboards are still led by general models.