VulnLLM-R-7B

source code · reasoning

A reasoning model for source-code vulnerability detection (C/C++, Python, Java), distilled from larger models over Qwen2.5-7B.

Context8,192
Curationcurated
AvailabilityAvailable
ReasoningThinking-enabled
Price / Mtok$0.15 in · $0.60 out
BaseQwen2.5-7B-Instruct
QuantizationQ4_K_M
LicenseApache-2.0

Source: Virtue-AI-HUB/VulnLLM-R-7B on HuggingFace

Good for
  • +source-code vulns
  • +C/C++/Python/Java
  • +reasoning
Model cardfrom the published weights — parameters, architecture, training, intended use
Parameters7B
ArchitectureQwen2.5 7B (Qwen2ForCausalLM)
Modalitytext
Training

Reasoning-distilled from larger models for function/project-level vuln detection over Qwen2.5-7B-Instruct.

Intended use
  • ›Reasoning-based source-code vulnerability detection
  • ›Generate a Chain-of-Thought over data flow, control flow, and
  • ›Detect complex logic vulnerabilities across C, C++, Python, and Java
  • ›Serve as an efficient 7B alternative to large general-purpose
Call it from your own codethe API is OpenAI-compatible — an existing client needs one line changed

Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AD_API_KEY"],
    base_url="https://adversariallm.ai/v1",
)

stream = client.chat.completions.create(
    model="vulnllm-r-7b",
    messages=[{"role": "user", "content": "…"}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")
Full model cardcurated from the model's HuggingFace card — description, training, usage, limitations, reported benchmarks

VulnLLM-R is described by its authors as "the first specialized reasoning Large Language Model designed specifically for software vulnerability detection." Unlike traditional static-analysis tools (the card names CodeQL) or small LLMs that rely on simple pattern matching, it is trained to reason step-by-step about data flow, control flow, and security context, mimicking the thought process of a human security auditor to identify complex logic vulnerabilities. Rather than just classifying code, it generates a "Chain-of-Thought" that analyzes why a vulnerability exists before giving a final answer. The 7B model is built on Qwen2.5-7B-Instruct and is trained and tested on C, C++, Python, and Java, with the card claiming zero-shot generalization. The authors report state-of-the-art results on the PrimeVul, Juliet 1.3, and ARVO benchmarks (numeric metrics are deferred to the paper, arXiv:2512.07533).

Training data

The card does not list specific training datasets, corpus sizes, or a detailed training procedure. It states the model is built on the Qwen/Qwen2.5-7B-Instruct base and is "trained to reason step-by-step about data flow, control flow, and security context," and that it is "trained and tested on C, C++, Python, and Java (zero-shot generalization)." The benchmarks named for evaluation are PrimeVul, Juliet 1.3, and ARVO. The associated paper title ("VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection", arXiv:2512.07533) references an "Agent Scaffold," and full data/method details are pointed to the paper and GitHub repo (github.com/ucsb-mlsec/VulnLLM-R) rather than given on the card itself.

Intended use
  • +Reasoning-based source-code vulnerability detection: analyze a code snippet step-by-step to determine whether it contains a vulnerability and explain why
  • +Generate a Chain-of-Thought over data flow, control flow, and security context, emulating a human security auditor, rather than binary pattern-matching classification
  • +Detect complex logic vulnerabilities across C, C++, Python, and Java (authors claim zero-shot generalization to these languages)
  • +Serve as an efficient 7B alternative to large general-purpose reasoning models and static-analysis tools for vulnerability-detection research
Limitations
  • −The card contains NO explicit limitations or out-of-scope / safety section
  • −No numeric benchmark results are given on the card; all performance figures are deferred to the paper (Figure 1 and Table 4, arXiv:2512.07533)
  • −Maximum context length is not stated on the card (base Qwen2.5-7B-Instruct supports long context, but the card does not confirm any value)
  • −Language coverage is stated only for C, C++, Python, and Java; behavior on other languages is unspecified
  • −Superiority claims over Claude-3.7-Sonnet, o3-mini, CodeQL, and AFL++ are vendor-reported and not substantiated with numbers on the card
Running the weights yourself

The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_name = "UCSB-SURFI/VulnLLM-R-7B"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name, 
    torch_dtype=torch.bfloat16, 
    device_map="auto"
)

# Example Code Snippet
code_snippet = """
void vulnerable_function(char *input) {
    char buffer[50];
    strcpy(buffer, input); // Potential buffer overflow
}
"""

# Prompt Template (Triggering Reasoning)
prompt = f"""You are an advanced vulnerability detection model. 
Please analyze the following code step-by-step to determine if it contains a vulnerability.

Code:
{code_snippet}

Please provide your reasoning followed by the final answer.
"""

messages = [
    {"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

generated_ids = model.generate(
    model_inputs.input_ids,
    max_new_tokens=512
)
generated_ids = [
    output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
]

response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(response)
Benchmarks — as reported on the card

Reported by the model’s authors, not our own testing — our scores are in the table above.

PrimeVulnot reported on cardvendor-reported; named as a benchmark where authors claim SOTA, but no numeric metric is given on the card (deferred to paper Fig 1 / Table 4, arXiv:2512.07533)
Juliet 1.3not reported on cardvendor-reported; named benchmark, no numeric result on the card (see paper Table 4)
ARVOnot reported on cardvendor-reported; named benchmark, no numeric result on the card (see paper Table 4)
Scores

No measurements published for this version yet.

Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.

Versionsscores attach to a version; v2 does not inherit v1's numbers
v12026-08-07—current

Start with VulnLLM-R-7B

A confirmed email account includes 30 free messages a month.

Sign in to start
VulnLLM-R-7B — Vulnerability detection · AdversariaLLM