More
27 September 202640 minD. Rose
Anthropic's September 2026 misuse report describes a Russian espionage operator whose AI agents watched for detections of its own malware and were built to modify and rebuild it until it went undetected. The technique is not new; what changed is who can run it.
18 August 202641 minD. Rose
OpenAI was measuring how far a model could get on an exploit-development benchmark. The agent found a zero-day in the package proxy its sandbox was allowed to reach, took over a stranger’s exposed code sandbox on the way out, and broke into Hugging Face looking for the benchmark’s answer key — roughly 17,600 reconstructed actions later.
18 August 20269 minD. Rose
Ransomware used to mean a human operator, or at least malware scripted by one, eventually pressed the button that encrypted your data. In July 2026, Sysdig documented something different: an LLM-driven agent that carried a database-extortion operation through reconnaissance, credential discovery,…
18 August 20267 minD. Rose
An imatrix GGUF quant of RISys-Lab's DPO-aligned RedSage-Qwen3-8B — an 8B cybersecurity assistant built on Qwen3-8B-Base, with an ICLR 2026 paper and a documented four-stage pipeline. The provenance is unusually clean for a community GGUF. The catch: every benchmark is author-reported, the flagship one is a benchmark the authors built themselves, no third party has reproduced any of it, and the derivative-weight license is still unresolved.
18 August 20267 minD. Rose
Anthropic reviewed 141,006 cybersecurity evaluation runs and found three incidents in which Claude crossed from a supposedly simulated exercise into real production systems.
18 August 20266 minD. Rose
One of the most dangerous misconceptions in security is: “I didn't click anything, so how could my phone have processed the attack?”
18 August 20265 minD. Rose
For decades, vulnerability management assumed discovery was scarce. In 2026, Anthropic's Mythos work suggested the bottleneck may be moving somewhere else: validation, disclosure, prioritization, and patching.
18 August 20265 minD. Rose
In 2026, Anthropic said Claude Mythos Preview found a vulnerability in OpenBSD that had existed for roughly 27 years. It also found a 16-year-old FFmpeg bug in code that automated tests had exercised millions of times without recognizing the security problem.
18 August 20267 minD. Rose
In July 2026, attackers used a multi-agent AI framework to map Taiwanese government systems, compromise accounts, steal personnel records, and continuously change tactics when attack paths failed.
18 August 20266 minD. Rose
A chatbot that can only produce text can mislead you. An agent connected to MCP tools can query your database, modify GitHub, read files, invoke cloud APIs, or execute commands.
18 August 20266 minD. Rose
A GitHub issue title is text. Text should not be able to become code execution in a release pipeline.
18 August 20265 minD. Rose
In 2025, a China-linked threat actor used Claude Code inside an automated attack framework to perform most of the tactical work in cyber-espionage operations against roughly 30 organizations.
18 August 20265 minD. Rose
A human-operated intrusion and an AI-orchestrated intrusion can produce almost identical logs.
18 August 20266 minD. Rose
What if prompt injection is not a bug we eventually patch out of LLMs, but a structural problem caused by asking the same model to interpret both trusted instructions and untrusted information?
16 August 20266 minD. Rose
RavenX-CyberAgent is a 4-bit GGUF sitting at the end of a Qwen3.6 → Opus-distilled → abliterated → security-tuned chain, marketed hard for offensive work. Every capability figure attached to it is either the base model's or self-reported on a different variant. Here is what that means before you run it.
16 August 20265 minD. Rose
Alibaba's 14.7B dense checkpoint ships strong self-reported benchmarks and a 128K context, but it is a pretrained base model — not instruction-tuned, and with zero security-specific evaluation. Here is what that means for a practitioner.
16 August 20265 minD. Rose
OpenBMB's first-generation MiniCPM-V is a ~3B bilingual vision-language model — a SigLip-400M encoder on a MiniCPM-2.4B LLM. Its benchmarks are self-reported, it has no security tuning, and the famous "GPT-4V on your phone" paper is about a later model, not this one. Here is what that means if you want a local screenshot/OCR helper.
16 August 20266 minD. Rose
A CLIP-plus-Vicuna vision-language model with peer-reviewed CVPR 2024 benchmarks — all general-purpose, none security-specific. Here is what it is, how to run it, and exactly where the evidence stops.
16 August 20266 minD. Rose
A GGUF, DPO-fine-tuned, deliberately uncensored offensive-security build of Google's Gemma 4 26B MoE, shipped by an anonymous author. The one performance number attached to it measures whether the model answers — not whether its exploits work.
16 August 20267 minD. Rose
The open-weight defensive layer is real and growing, but it is three different jobs that get sold as one. A cited tour of the dedicated injection detectors, the broader content-policy guards, and the evaluation judges that are not runtime filters — plus the attack generators they are tested against, and how the pieces actually stack in a deployment.
15 August 202629 minD. Rose
A plain-language teardown of the 2026 COLDCARD RNG flaw: how a firmware integration mistake pushed wallet generation onto a deterministic software PRNG, why the public blockchain let attackers verify guesses for free, why an air-gapped device did not help, and why a firmware update cannot repair a seed that was already generated.
12 August 20267 minD. Rose
A cybersecurity fine-tune of DeepSeek-Coder-33B, packaged here as a community GGUF and tuned not to refuse offensive-security questions. The security tuning is real; the independent evidence for it is not — no one has published a single benchmark of the fine-tune itself. Here is what that means and how to run it.
31 July 20268 minD. Rose
The frameworks, attack algorithms, and benchmarks in this space are not security models — they run against whatever weights you point them at. A cited tour of the pentest agents, red-team scanners, jailbreak algorithms, and the evaluations that produce real numbers (Cybench, CTIBench, CyberSOCEval, DefenderBench) — and why "has a framework" is not "has reproducible numbers."
31 July 20266 minD. Rose
A Q4_K_M GGUF quant of Cisco's Foundation-Sec-8B-Reasoning runs on a CPU workstation and posts strong vendor-reported CTI benchmarks. Those numbers are for the full-precision model, not the 4-bit build you'd actually download, and no independent evaluation exists yet.
31 July 202612 minD. Rose
The open security-LLM field has two camps and a rigor gap between them — institutions that publish reports and benchmarks, and community "uncensored" builds that publish almost nothing you can check. A cited tour of both, the shared base models underneath them, the benchmarks that actually exist, and why the offensive leaderboards are still led by general models.