Qwen 2.5 14B
Fast general-purpose
General-purpose fast helper model.
Alibaba's Qwen2.5-14B-Instruct: pretrained on ~18T tokens, then instruction-tuned for chat, reasoning, long context and code. A strong general model that follows instructions cleanly — the catalog's most reliable default.
- ›General security research and Q&A that isn't tied to one specialist model
- ›Writing and explaining detection rules, queries and scripts
- ›Structured output — tables, JSON, step-by-step analysis
- ›Long-context work up to 32K tokens
Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AD_API_KEY"],
base_url="https://adversariallm.ai/v1",
)
stream = client.chat.completions.create(
model="qwen25-14b",
messages=[{"role": "user", "content": "…"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")Qwen2.5-14B-Instruct is the instruction-tuned 14B model in the Qwen2.5 series from Alibaba Cloud, a family of base and instruction-tuned LLMs ranging from 0.5B to 72B parameters. The card describes it as a causal language model built on the transformer architecture with RoPE, SwiGLU, RMSNorm, and Attention QKV bias. It has 14.7B total parameters (13.1B non-embedding), 48 layers, and grouped-query attention with 40 query heads and 8 key/value heads. The card highlights improvements over prior Qwen releases including more knowledge, stronger coding and mathematics ability, better instruction following, long-text generation over 8K tokens, and improved handling of structured data. It supports a context length of up to 131,072 (128K) tokens, generation up to 8,192 tokens, and multilingual use across more than 29 languages.
The card only classifies the training stage as "Pretraining & Post-training." It does not disclose training datasets, data composition, or the specific fine-tuning methodology used for the instruction-tuned model.
- +Instruction-tuned conversational assistant used via the transformers chat template (apply_chat_template) with system/user/assistant roles
- +Coding and mathematics tasks (card cites improved coding and math capabilities)
- +Long-form text generation exceeding 8K tokens
- +Structured-data tasks such as understanding tables and generating structured outputs (e.g. JSON)
- +Long-context workloads up to 128K tokens, with YaRN rope_scaling enabled when inputs exceed 32,768 tokens
- +Multilingual applications across 29+ languages
- −Requires transformers >= 4.37.0; older versions raise a KeyError for the qwen2 model type
- −Long-context (>32,768 tokens) requires manually enabling YaRN rope_scaling in config; the card advises adding it only when long inputs are actually needed
- −vLLM only supports static YaRN, so the scaling factor stays constant regardless of input length, which the card notes may impact performance on shorter texts
- −The card does not document safety, bias, or out-of-scope-use guidance beyond the above technical notes
The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "Qwen/Qwen2.5-14B-Instruct"
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype="auto",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
prompt = "Give me a short introduction to large language model."
messages = [
{"role": "system", "content": "You are Qwen, created by Alibaba Cloud. You are a helpful assistant."},
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
generated_ids = model.generate(
**model_inputs,
max_new_tokens=512
)
generated_ids = [
output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
]
response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]No measurements published for this version yet.
Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.
Alibaba's 14.7B dense checkpoint ships strong self-reported benchmarks and a 128K context, but it is a pretrained base model — not instruction-tuned, and with zero security-specific evaluation. Here is what that means for a practitioner.
Read the analysis →Start with Qwen 2.5 14B
A confirmed email account includes 30 free messages a month.