Llama-Primus-Reasoning
CVE/CWE · security reasoning
Trend Micro's reasoning model, continued-pretrained on the open Primus cybersecurity corpus over Llama-3.1-8B.
Source: trendmicro-ailab/Llama-Primus-Reasoning on HuggingFace
- +CTI
- +CVE/CWE/CVSS
- +reasoning
Continued pretraining + reasoning tuning on the open Primus cyber datasets over Llama-3.1-8B.
- ›Cybersecurity reasoning / chain-of-thought over security tasks (the
- ›Security-certification-style question answering — the card's only
- ›Research use in cybersecurity LLM work; the gated access form lists
- ›A lightweight (8B) open model for experimenting within Trend Micro's
Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AD_API_KEY"],
base_url="https://adversariallm.ai/v1",
)
stream = client.chat.completions.create(
model="llama-primus-reasoning",
messages=[{"role": "user", "content": "…"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")Llama-Primus-Reasoning is, per Trend Micro, the "first cybersecurity reasoning model": a chain-of-thought model distilled from the reasoning steps and reflection data that o1-preview and DeepSeek-R1 generated on cybersecurity tasks (the Primus-Reasoning dataset), built on top of Llama-Primus-Merged. It is one model in the Primus collection, an open family of lightweight cybersecurity models and datasets spanning pre-training (Primus-Seed, Primus-FineWeb), instruction fine-tuning (Primus-Instruct) and reasoning distillation (Primus-Reasoning), all layered on Llama-3.1-8B-Instruct. On the CISSP security-certification benchmark the card reports a 15.8% improvement over Llama-3.1-8B-Instruct. Per the 2025/06/02 news note, the released checkpoint was re-distilled jointly from DeepSeek-R1 and o1-preview after Primus-Reasoning was expanded with additional DeepSeek-R1 samples, which the vendor reports as its best-performing CISSP variant. The card states the datasets and models share the research foundation behind Trend Micro's Trend Cybertron product, and that no Trend Micro customer information is included.
Reasoning-distillation fine-tune of Llama-Primus-Merged on Primus-Reasoning: reasoning-step + reflection traces generated by o1-preview and DeepSeek-R1 on cybersecurity tasks. The 2025/06/02 update expanded Primus-Reasoning with additional DeepSeek-R1 samples and re-distilled the model jointly from DeepSeek-R1 + o1-preview (the released version). The broader Primus collection it draws lineage from also includes Primus-Seed and Primus-FineWeb (cybersecurity pre-training) and Primus-Instruct (instruction fine-tuning), all on a Llama-3.1-8B-Instruct base. Frontmatter datasets: Primus-Reasoning, Primus-Seed, Primus-FineWeb, Primus-Instruct. The card explicitly notes: "No TrendMicro customer information is included." (Sample sizes seen on the HF widget — Primus-Seed ~174k, Primus-FineWeb ~3.39M, Primus-Instruct ~835 — come from the linked dataset cards, not this README, so treat as unconfirmed here.)
- +Cybersecurity reasoning / chain-of-thought over security tasks (the model's stated purpose as a 'cybersecurity reasoning model')
- +Security-certification-style question answering — the card's only evaluation is CISSP (Certified Information Systems Security Professional)
- +Research use in cybersecurity LLM work; the gated access form lists Research / Commercial / Other as intended-use options
- +A lightweight (8B) open model for experimenting within Trend Micro's Primus collection of cyber datasets and models
- −Gated model: access requires accepting terms and submitting Affiliation, Country, intended-use, Job title, plus IP geolocation.
- −Evaluated only on CISSP; the card reports no other cybersecurity, coding, general-reasoning, or safety benchmarks.
- −MIT-based but use additionally requires compliance with the Llama 3.1 Community License Agreement.
- −The reasoning variant produces long outputs — ~1,467 avg tokens per CISSP answer vs ~1 for the no-CoT 5-shot baseline — increasing latency/cost.
- −Inherits Llama-3.1-8B limitations, though the card does not enumerate them.
- −No explicit out-of-scope, safety, or bias section is provided; the only caveat stated is that no Trend Micro customer data is included.
The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.
The model card provides NO code or prompt usage example (no Transformers/vLLM/SGLang snippet, no sample prompt, no BibTeX in the README body). Metadata only: library_name: transformers, pipeline_tag: text-generation, base_model: trendmicro-ailab/Llama-Primus-Merged. Paper: arXiv:2502.11191.
Reported by the model’s authors, not our own testing — our scores are in the table above.
| CISSP (w/ CoT, 0-shot) — Llama-Primus-Reasoning [Llama-Primus-Merged + distilled from o1-preview + DeepSeek-R1] | 0.8193 | This IS the released model / best variant; ↑15.8% over the Llama-3.1-8B-Instruct w/o-CoT 5-shot baseline (0.7073); avg 1467.40 tokens. Vendor-reported (Trend Micro, arXiv:2502.11191). |
| CISSP (w/o CoT, 5-shot) — Llama-3.1-8B-Instruct (baseline) | 0.7073 | Reference baseline for the ↑% column; avg 1 token. Vendor-reported. |
| CISSP (w/ CoT, 0-shot) — Llama-Primus-Merged (no reasoning distillation) | 0.7603 | ↑7.49%; avg 241.92 tokens. Vendor-reported; shows the pre-distillation baseline of this model. |
| CISSP (w/ CoT, 0-shot) — Llama-Primus-Merged + distilled from DeepSeek-R1 only | 0.8075 | ↑14.2%; avg 1483.94 tokens. Vendor-reported. |
| CISSP (w/ CoT, 0-shot) — Llama-Primus-Merged + distilled from o1-preview only | 0.7780 | ↑10.0%; avg 726.96 tokens. Vendor-reported. |
| CISSP — o1-preview (raw model, comparison) | 0.8035 | avg 1054.91 tokens. Vendor-reported comparison row. |
| CISSP — DeepSeek-R1 (raw model, comparison) | 0.8212 | avg 1229.32 tokens. Vendor-reported; note full DeepSeek-R1 slightly outscores this 8B model (0.8212 vs 0.8193). |
| CISSP — DeepSeek-R1-Distill-Llama-8B (comparison) | 0.7399 | ↑4.61%; avg 1542.10 tokens. Vendor-reported comparison; this Primus model outscores it. |
No measurements published for this version yet.
Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.
Start with Llama-Primus-Reasoning
A confirmed email account includes 30 free messages a month.