Foundation-Sec-8B-Reasoning
Reasoning
Security reasoning helper model.
Source: fdtn-ai/Foundation-Sec-8B-Reasoning-Q4_K_M-GGUF on HuggingFace
Fine-tuned by Foundation AI from Meta Llama-3.1-8B on a curated corpus of cybersecurity text — CVE and CWE records, threat intelligence, exploit and detection writeups — to reason over security tasks rather than to chat.
- ›Reasoning about a vulnerability's exploitability from an advisory
- ›Explaining a CWE class and how it manifests in code
- ›Drafting detection logic from a described technique
- ›Summarising a threat-intel report into IOCs and TTPs
Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AD_API_KEY"],
base_url="https://adversariallm.ai/v1",
)
stream = client.chat.completions.create(
model="foundation-sec-8b",
messages=[{"role": "user", "content": "…"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")Foundation-Sec-8B-Reasoning (full name Llama-3.1-FoundationAI-SecurityLLM-8B-Reasoning) is an open-weight, 8-billion-parameter instruction-tuned language model developed by Foundation AI at Cisco and specialized for cybersecurity applications. Its architecture is an auto-regressive transformer built on a Meta Llama-3.1-8B backbone, and it extends the Foundation-Sec-8B base model with instruction-following and reasoning capabilities. Released January 28, 2026, it targets security workflows such as SOC triage, threat-intelligence reasoning, vulnerability analysis, and incident enrichment. Its training-data cutoff is April 10, 2025, and it was trained with a 32,768-token sequence length using the AdamW optimizer. It supports summarization, classification, named-entity recognition, question answering, and multi-step reasoning across security domains.
Extends the Foundation-Sec-8B base model. Per the card, Foundation-Sec-8B-Reasoning "was trained on a wide variety of public and proprietary question answer/pairs for general and security-specific reasoning and instruction-following tasks" (data cutoff April 10, 2025; 32,768-token sequence length; AdamW optimizer). The underlying Foundation-Sec-8B base was continued-pretrained on approximately 5.1 billion tokens of cybersecurity-specific data — threat-intelligence reports, vulnerability databases, and incident-response documentation — described as "meticulously collected from public sources on the web" via large-scale web crawling, relevancy filtering, deduplication, and quality filtering (4096-token sequence length for the base).
- +SOC Acceleration: automating alert triage, summarization, case documentation, and evidence collection
- +Proactive Threat Defense: simulating attacks and modeling attacker behavior
- +Engineering Enablement: providing security assistance and compliance/product-security assessment
- +Downstream tasks: summarization of incident reports, threat classification, named-entity extraction, and Q&A for SOC analysts
- +Reasoning tasks across threat mapping, vulnerability prioritization, and incident enrichment
- −No knowledge of vulnerabilities, threats, or best practices after the April 2025 training cutoff
- −May reflect biases present in the security literature it was trained on
- −Cannot verify the identity or legitimate intent of the user
- −Struggles with complex multi-step security scenarios
- −Cannot access external systems or independently verify factual accuracy (can hallucinate)
- −Out-of-scope: generating malware / phishing / exploitation techniques or other harmful content, autonomous security decisions without human oversight, legal or medical advice, and general non-security tasks or any unlawful use
- −Vendor deployment guidance: keep a human in the loop, add validation layers, use structured prompting emphasizing ethical practices, and supplement with current threat feeds via RAG
The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("fdtn-ai/Foundation-Sec-8B-Reasoning")
model = AutoModelForCausalLM.from_pretrained("fdtn-ai/Foundation-Sec-8B-Reasoning")
prompt = "CVE-2015-10011 is a vulnerability about OpenDNS OpenResolve improper log output neutralization. What is the corresponding CWE?"
messages = [{"role": "user", "content": prompt}]
model_inputs = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(model_inputs, return_tensors="pt", add_special_tokens=False)
output = model.generate(**inputs, temperature=0.1, max_new_tokens=1024)
resp = tokenizer.batch_decode(output)[0]
print(resp.replace(model_inputs, ""))Reported by the model’s authors, not our own testing — our scores are in the table above.
| CTI-MCQA (threat-intel multiple-choice QA) | 0.691 (vs Llama 3.1 8B 0.607, GPT-5-Nano 0.688) | vendor-reported (Cisco Foundation AI eval suite), not independently verified |
| CTI-RCM (root-cause mapping) | 0.753 (vs Llama 3.1 8B 0.531, GPT-5-Nano 0.672) | vendor-reported (Cisco Foundation AI eval suite), not independently verified |
| CTI-VSP (vulnerability severity prediction) | 0.856 (vs Llama 3.1 8B 0.811, GPT-5-Nano 0.822) | vendor-reported (Cisco Foundation AI eval suite), not independently verified |
| CTI-Reasoning | 0.411 (vs Llama 3.1 8B 0.335, GPT-5-Nano 0.431 — competitor scores higher here) | vendor-reported (Cisco Foundation AI eval suite), not independently verified |
| HarmBench (safety refusal) | 93.00%, rising to 98.25% with LlamaGuard integration | vendor-reported safety figure, not independently verified |
No measurements published for this version yet.
Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.
Cisco published numbers for Foundation-Sec-8B with a technical report — a solid B for the full-precision model. The artifact in this catalog is the Q4_K_M quantisation, and no one has measured what that quantisation costs, so the evidence is about a different artifact than the one you would run.
This grades the RECORD, not the model. A model rated E may be excellent — the claim is only that nobody has shown it.
A Q4_K_M GGUF quant of Cisco's Foundation-Sec-8B-Reasoning runs on a CPU workstation and posts strong vendor-reported CTI benchmarks. Those numbers are for the full-precision model, not the 4-bit build you'd actually download, and no independent evaluation exists yet.
Read the analysis →Start with Foundation-Sec-8B-Reasoning
A confirmed email account includes 30 free messages a month.