Foundation-Sec-8B-Reasoning at 4-bit: Cisco's open security-reasoning model, and the fine print on its CTI scores
D. Rose · 31 July 2026 · Updated 16 August 2026 · 6 min
A Q4_K_M GGUF quant of Cisco's Foundation-Sec-8B-Reasoning runs on a CPU workstation and posts strong vendor-reported CTI benchmarks. Those numbers are for the full-precision model, not the 4-bit build you'd actually download, and no independent evaluation exists yet.
Cisco's Foundation AI group shipped something the open ecosystem was short on: a security-tuned reasoning model with published cyber-threat-intelligence benchmarks and a technical report behind them. The entry in this catalog is one specific build of it — the Q4_K_M 4-bit GGUF quantization of Foundation-Sec-8B-Reasoning, the version you'd pull to run on a laptop or an air-gapped box.
That distinction matters for everything below. The benchmarks are Cisco's, they're for the full-precision model, and the thing you'd actually download is a lossy 4-bit derivative of it that nobody has separately measured. Here's what's real, what's claimed, and what you'd use it for.
What it actually is
The full model is an 8-billion-parameter autoregressive transformer built on Meta Llama-3.1-8B, released in BF16 with a 32,768-token context window (model card). It is not a single fine-tune of stock Llama. Per the technical report, Cisco started from their own Foundation-Sec-8B base — itself continued pretraining of Llama-3.1-8B — and added instruction-following plus explicit reasoning traces through a two-stage process: supervised fine-tuning, then reinforcement learning from verifiable rewards (RLVR), on proprietary data spanning cybersecurity analysis, instruction-following, and mathematical reasoning.
The provenance chain, end to end:
Llama-3.1-8B → Foundation-Sec-8B (continued pretraining) → Foundation-Sec-8B-Reasoning (SFT + RLVR) → Q4_K_M GGUF (llama.cpp quantization)
Two dates worth pinning: the training-data cutoff is April 10, 2025, and the reasoning model was released January 28, 2026 (model card). And one trap for anyone who used the earlier release: the original Foundation-Sec-8B base is documented at a 4,096-token context, while the Reasoning model is 32,768. They are not interchangeable — don't carry assumptions from one to the other.
The catalog entry itself is the Q4_K_M GGUF: roughly 4.92 GB on disk, ~4.94 GB at runtime, converted via llama.cpp. It's built by Foundation AI at Cisco (Hugging Face org fdtn-ai; the quant card lists blainen@cisco.com as contact).
Makers' claims vs. what's verified
Cisco positions Foundation-Sec-8B-Reasoning as the "first open-weight security reasoning model," trained to reflect how security practitioners actually reason, and deployable on-prem or in air-gapped environments (Cisco blog). The model card frames three use-case categories (model card):
- SOC acceleration — triage, summarization, case-note generation, evidence collection.
- Proactive threat defense — simulating attacks, prioritizing vulnerabilities, mapping TTPs, modeling attacker behavior.
- Engineering enablement — security assistance, validating configurations, assessing compliance evidence.
Those are the author's positioning claims, not independently verified capabilities. What separates this model from a typical community upload is that the claims come with numbers and a report — but every one of those numbers traces back to Cisco's own model card, blog, and technical report. No independent, third-party evaluation of Foundation-Sec-8B-Reasoning was found. So read the benchmarks below as vendor-reported, and read the use-case list as intent rather than proof.
Benchmarks
Cisco reports these against Llama-3.1-8B and GPT-5-Nano as comparators. Every score here is for the full-precision Foundation-Sec-8B-Reasoning model, not the 4-bit GGUF in this catalog — there are no published benchmarks for the quant.
| Benchmark | Foundation-Sec-8B-Reasoning | Llama-3.1-8B | GPT-5-Nano | Which model / precision |
|---|---|---|---|---|
| CTI-MCQA (threat-intel multiple-choice QA) | 0.691 | 0.607 | 0.688 | Full-precision Reasoning model |
| CTI-RCM (root-cause mapping) | 0.753 | 0.531 | 0.672 | Full-precision Reasoning model |
| CTI-VSP (vulnerability severity prediction) | 0.856 | 0.811 | 0.822 | Full-precision Reasoning model |
| CTI-Reasoning | 0.411 | 0.335 | 0.431 | Full-precision Reasoning model (GPT-5-Nano leads) |
| HarmBench (attack resistance) | 93.00% alone / 98.25% + LlamaGuard | — | — | Full-precision Reasoning model |
Sources: CTI-MCQA and HarmBench from the model card; CTI-RCM, CTI-VSP, and CTI-Reasoning from the Cisco blog. Worth stating plainly: on CTI-Reasoning, the hardest of the four, GPT-5-Nano edges it out (0.431 vs 0.411). The story is "competitive with a small frontier model on CTI, at 8B and open weights," not "beats everything."
A separate set of numbers exists for the base Foundation-Sec-8B — a different model generation, reported on a percentage (0–100) scale rather than the 0–1 scale above. Do not cross-compare the two tables.
| Benchmark (percentage scale) | Foundation-Sec-8B (base) | Llama-3.1-8B | Llama-3.1-70B |
|---|---|---|---|
| CTI-MCQA | 67.39 | 64.14 | 68.23 |
| CTI-RCM | 75.26 | 66.43 | 72.66 |
Source: base model card. One honest caveat on the scales themselves: the 0–1 vs. 0–100 reading is inferred from how the two cards present their tables, not confirmed against a benchmark spec. The technical report states evaluation across 10 cybersecurity and 10 general-purpose benchmarks; only the four CTI tasks plus HarmBench surfaced in the card and blog, so the full table — including general-purpose scores — isn't reflected here.
For a practitioner
What you'd genuinely use it for. CTI triage that fits on hardware you control. The reported strengths — root-cause mapping (CTI-RCM 0.753) and vulnerability severity prediction (CTI-VSP 0.856) — line up with real SOC and vuln-management drudgery: turning a CVE or an alert into a structured severity call or a mapped root cause. The pitch that earns attention here is on-prem / air-gapped deployment (Cisco blog): sensitive incident data never leaves the box.
How to run it. The Q4_K_M GGUF targets local and CPU inference through llama.cpp (and is Ollama-compatible). At ~4.92 GB on disk and ~4.94 GB of runtime memory, it runs on a modest workstation with no GPU. The full model's context window is 32K; note the quant card doesn't restate it, so confirm your context settings rather than assuming. Cisco publishes a usage cookbook with examples.
The caveats that will actually bite you:
- It's a reasoning model. It emits explicit chain-of-thought before answers, so outputs are longer and token usage is higher than the plain base/instruct variants. Budget for that latency in any SOC pipeline where response time matters.
- You're running a lossy 4-bit build of a model that was benchmarked at full precision. The CTI scores above will degrade at Q4 by an unmeasured amount — Cisco publishes no quant-specific numbers. If accuracy matters more than footprint, higher-fidelity quants (e.g. a
Q8_0GGUF) exist in thefdtn-aiorg. - It is not meant to be your only safety layer. Cisco's best HarmBench result (98.25%) is the model paired with an external LlamaGuard filter; standalone it's 93.00% (model card). The framing implies a filter in front of it, not the model alone.
Licensing — read before you ship. The base Foundation-Sec-8B is cleanly Apache 2.0 (base card), and the GGUF quant card claims "Apache 2.0 (same as base)" (quant card). But the upstream Reasoning card lists its license as "other" and points to a NOTICE.md rather than a plain Apache-2.0 tag (model card). Read that NOTICE directly before relying on Apache terms — and remember the derivative also inherits Meta Llama 3.1 licensing obligations upstream. This one isn't fully resolved.
Honest limits
- No benchmarks exist for the 4-bit quant — the exact build in this catalog. Every score is from the full-precision model; quantization loss is unmeasured.
- No independent evaluation. All CTI numbers trace to Cisco's own card, blog, and report. Treat them as vendor-reported until a third party reproduces them.
- The quant card itself is thin — it publishes no context length, no benchmarks, and no quant-quality metrics. Everything about the underlying model here is inferred from the upstream full-precision card and report.
- The benchmark scale is inferred. Whether the CTI tasks are 0–1 accuracy or 0–100 percentages is read off the tables, not confirmed against a spec — which is exactly why the two tables above are kept separate.
Foundation-Sec-8B-Reasoning is one of the more credible open security models to look at right now: real report, real CTI numbers, an 8B footprint that fits on a workstation, and a maker who put a name and an email on it. The discipline is keeping the labels straight — full-precision vs. 4-bit, Reasoning model vs. base, vendor-reported vs. independent — which is the whole reason it lives in a measured catalog.
See the catalog entry: /catalog/foundation-sec-8b. For where it sits among the rest of the field, see the 2026 survey of open security-research LLMs.