CyberPal 2.0 20B
threat correlation · SOC
IBM Research's security model fine-tuned on enriched cyber chain-of-thought data (SecKnowledge 2.0) over a gpt-oss-20B MoE base.
Source: cyber-pal-security/CyberPal2.0-20B on HuggingFace
- +CTI knowledge
- +threat investigation
- +vuln correlation
Fine-tuned on SecKnowledge 2.0 chain-of-thought cyber data over gpt-oss-20b (IBM Research, arXiv:2510.14113).
- ›Cyber Threat Intelligence (CTI)
- ›Vulnerability & weakness analysis
- ›Detection & mitigation guidance
- ›Security operations support
Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AD_API_KEY"],
base_url="https://adversariallm.ai/v1",
)
stream = client.chat.completions.create(
model="cyberpal-2-20b",
messages=[{"role": "user", "content": "…"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")CyberPal-2.0-20B is a cybersecurity-expert 20B-parameter Small Language Model (SLM) fine-tuned from gpt-oss-20b for security operations and threat-management workflows — e.g., cyber threat intelligence (CTI) Q&A, vulnerability-to-weakness mapping, and detection/mitigation recommendations. It is part of the CyberPal 2.0 model family (4B–20B) and is tied to the paper "Toward Cybersecurity-Expert Small Language Models" (arXiv:2510.14113, IBM Research; authors Levi, Ohayon, Blobstein, Sagi, Molloy, Allouche). The card describes it as a decoder-only, instruction-tuned model trained with the SecKnowledge 2.0 data enrichment and formatting pipeline to produce higher-fidelity, task-grounded reasoning traces (chain-of-thought) for cybersecurity tasks. Its stated domain focus spans CTI, SOC/IR, appsec, IAM, and governance/compliance. The card lists an 8192-token context length, BF16 tensors, and Apache-2.0 licensing, and explicitly frames outputs as advisory rather than authoritative.
Fine-tuned on SecKnowledge 2.0, described by the card as the output of an enrichment + formatting pipeline that: (1) uses expert-in-the-loop schema/format steering (task-specific reasoning formats), (2) performs multi-step grounding using documents and/or web search, and (3) uses LLM-based judging for readability/factuality checks. Per the card, the starting SecKnowledge dataset comprises ~153k instructions in the first stage (from structured public security sources), expanded to a ~403k-example cybersecurity corpus in the second stage via synthetic generation. The SecKnowledge 2.0 pipeline is stated to use gpt-oss-120b (Medium reasoning effort) as the backbone LLM for dataset generation/enrichment. No raw pretraining corpus or full data-source list is enumerated on the card beyond "structured public security sources."
- +Cyber Threat Intelligence (CTI): answering CTI questions; mapping campaigns/actors/techniques; explaining ATT&CK concepts
- +Vulnerability & weakness analysis: correlating CVE evidence / bug tickets to CWE root causes
- +Detection & mitigation guidance: proposing detections/mitigations for tactics/techniques/weaknesses/vulnerabilities
- +Security operations support: incident summarization, investigation assistance, hypothesis-driven triage, response recommendations
- −Out-of-scope: any form of wrongdoing, intrusion, malware development, or instructions intended to enable harm (card explicitly disallows)
- −Not recommended for high-stakes decisions without human review — the card says to treat outputs as advisory, not authoritative
- −Agentic / tool use is currently not tested; the card says support will be released in newer versions
- −No benchmark or evaluation results are provided on the model card itself (performance claims live in the paper, not the card)
- −No chat template / system-prompt format is documented on the card; usage examples pass a raw prompt string with no chat template applied
The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.
```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "cyber-pal-security/CyberPal-2.0-20B"
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
prompt = """Analyze this C code snippet. Identify the vulnerability, the potential impact, and provide a patched version.
void process_data(char *input) {
char buffer[128];
strcpy(buffer, input);
printf("Processed: %s", buffer);
}"""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
out = model.generate(
**inputs,
max_new_tokens=512,
do_sample=False,
temperature=0.0,
)
print(tokenizer.decode(out[0], skip_special_tokens=True))
```
The card also provides a vLLM example (LLM(model=model_id, dtype="bfloat16", tensor_parallel_size=1); SamplingParams(max_tokens=512, temperature=0.0)) and notes that newer vLLM versions should set `export VLLM_USE_FLASHINFER_MOE_FP16=0` to avoid numerical issues.No measurements published for this version yet.
Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.
Start with CyberPal 2.0 20B
A confirmed email account includes 30 free messages a month.