Llama-Primus-Merged
CTI · CISSP · CVE mapping
Trend Micro's Primus continued-pretrain merged into Llama-3.1-8B-Instruct — domain gains with general capability retained.
Source: trendmicro-ailab/Llama-Primus-Merged on HuggingFace
- +CTI
- +CISSP-style
- +CVE mapping
Primus cyber continued-pretrain merged into Llama-3.1-8B-Instruct (Trend Micro).
- ›Cybersecurity question-answering and knowledge tasks (CISSP-style
- ›Cyber threat intelligence (CTI) tasks
- ›Security-domain assistant use while retaining general
- ›Bilingual (English + Japanese) text-generation in the security domain
- ›Research use under gated access (card requires affiliation, country,
Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AD_API_KEY"],
base_url="https://adversariallm.ai/v1",
)
stream = client.chat.completions.create(
model="llama-primus-merged",
messages=[{"role": "user", "content": "…"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")Llama-Primus-Merged is a cybersecurity-specialized 8B LLM from Trend Micro's AI Lab, built on Llama-3.1-8B-Instruct. Per the card, it was first pre-trained on a large cybersecurity corpus (2.77B tokens), then instruction fine-tuned on approximately 1,000 curated cybersecurity QA tasks, and finally merged back with Llama-3.1-8B-Instruct to maintain the same instruction-following capability. The card reports that this merge preserves general chat ability while achieving a 14.84% improvement in aggregated scores across cybersecurity benchmarks versus the base Llama-3.1-8B-Instruct. The intro frames it in the context of domain-specialized LLMs ("Large Language Models (LLMs) have demonstrated remarkable versatility in recent years, with promising applications in specialized domains such as finance, law, and biomedicine."). The card carries languages English and Japanese and a text-generation pipeline tag, and explicitly notes: "No TrendMicro customer information is included."
Three-stage pipeline over the Primus dataset collection: (1) continued pre-training on ~2.77B tokens of cybersecurity text from Primus-Seed and Primus-FineWeb (producing Llama-Primus-Base); (2) instruction fine-tuning on ~1,000 curated cybersecurity QA tasks (Primus-Instruct); (3) model merging with Llama-3.1-8B-Instruct to retain instruction-following ability. YAML-listed datasets: trendmicro-ailab/Primus-Seed, trendmicro-ailab/Primus-FineWeb, trendmicro-ailab/Primus-Instruct.
- +Cybersecurity question-answering and knowledge tasks (CISSP-style exams, SecEval, CyberMetric)
- +Cyber threat intelligence (CTI) tasks: MCQ, CVE-to-CWE mapping, CVSS scoring, ATE — per CTI-Bench evaluation
- +Security-domain assistant use while retaining general instruction-following / chat capability
- +Bilingual (English + Japanese) text-generation in the security domain
- +Research use under gated access (card requires affiliation, country, intended-use, job title, and geolocation verification)
- −Multilingual MMLU slightly regresses vs base: English 67.36 vs 68.16, Japanese 47.85 vs 49.22, French 58.14 vs 58.91, German 56.68 vs 57.70
- −Long-context (LongBench) slightly lower: 8K+ 50.66 vs 51.08, 16K+ 27.13 vs 29.18
- −General chat (MT Bench) marginally lower: 8.29375 vs 8.3491
- −Safety trade-offs reported: DAN jailbreak resistance drops (41.70% vs 28.98% — higher attack success), malwaregen (disallowed) rises to 29.00% vs 14.34%, XSTest over-refusal 83.20% vs 93.20%; some other safety metrics improve (e.g. snowball hallucination)
- −Governed by both MIT and the Llama 3.1 Community License Agreement, and gated/geolocation-verified access on HuggingFace
- −Max context length is not stated on the card
The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.
from transformers import pipeline
pipe = pipeline("text-generation", model="trendmicro-ailab/Llama-Primus-Merged")
messages = [
{"role": "user", "content": "Who are you?"},
]
pipe(messages)
```
```python
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("trendmicro-ailab/Llama-Primus-Merged")
model = AutoModelForCausalLM.from_pretrained("trendmicro-ailab/Llama-Primus-Merged", device_map="auto")
messages = [
{"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))Reported by the model’s authors, not our own testing — our scores are in the table above.
| CTI-Bench (MCQ) | 0.6656 | vendor-reported (Trend Micro); base Llama-3.1-8B-Instruct 0.6420 (higher better) |
| CTI-Bench (CVE -> CWE) | 0.6620 | vendor-reported; base 0.5910 (higher better) |
| CTI-Bench (CVSS) | 1.1233 | vendor-reported; base 1.2712 (lower better) |
| CTI-Bench (ATE) | 0.3387 | vendor-reported; base 0.2721 (higher better) |
| CyberMetric (500) | 0.8660 | vendor-reported; base 0.8560 |
| SecEval | 0.5062 | vendor-reported; base 0.4966 |
| CISSP (exams in book) | 0.7191 | vendor-reported; base 0.7073 |
| Cybersecurity Aggregate | 2.63 | vendor-reported; base 2.29; card claims +14.84% aggregate improvement |
| BFCL (V2) Function Calling | 74.77% | vendor-reported; base 73.02% |
| MT Bench (general chat) | 8.29375 | vendor-reported; base 8.3491 (slightly lower) |
| LongBench 8K+ | 50.66 | vendor-reported; base 51.08 (slightly lower) |
| LongBench 16K+ | 27.13 | vendor-reported; base 29.18 (slightly lower) |
| MMLU English | 67.36% | vendor-reported; base 68.16% |
| MMLU Japanese | 47.85% | vendor-reported; base 49.22% |
| MMLU French | 58.14% | vendor-reported; base 58.91% |
| MMLU German | 56.68% | vendor-reported; base 57.70% |
| garak DAN (Jailbreak resist) | 41.70% | vendor-reported; base 28.98% (higher = more resistant per card table) |
| garak malwaregen (Disallowed) | 29.00% | vendor-reported; base 14.34% |
| garak realtoxicityprompts | 85.40% | vendor-reported; base 90.03% |
| garak XSTest (Over-Refuse) | 83.20% | vendor-reported; base 93.20% |
No measurements published for this version yet.
Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.
Start with Llama-Primus-Merged
A confirmed email account includes 30 free messages a month.