Qwen 2.5 14B

Fast general-purpose

General-purpose fast helper model.

Context32,768
Curationcurated
AvailabilityAvailable
Price / Mtok$0.15 in · $0.60 out
BaseQwen2.5-14B
Quantization—
LicenseApache-2.0
Good for
    Model cardfrom the published weights — parameters, architecture, training, intended use
    Parameters14.7B
    ArchitectureQwen2.5 (Qwen2ForCausalLM)
    Modalitytext
    Training

    Alibaba's Qwen2.5-14B-Instruct: pretrained on ~18T tokens, then instruction-tuned for chat, reasoning, long context and code. A strong general model that follows instructions cleanly — the catalog's most reliable default.

    Intended use
    • ›General security research and Q&A that isn't tied to one specialist model
    • ›Writing and explaining detection rules, queries and scripts
    • ›Structured output — tables, JSON, step-by-step analysis
    • ›Long-context work up to 32K tokens
    Call it from your own codethe API is OpenAI-compatible — an existing client needs one line changed

    Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.

    import os
    from openai import OpenAI
    
    client = OpenAI(
        api_key=os.environ["AD_API_KEY"],
        base_url="https://adversariallm.ai/v1",
    )
    
    stream = client.chat.completions.create(
        model="qwen25-14b",
        messages=[{"role": "user", "content": "…"}],
        stream=True,
    )
    for chunk in stream:
        print(chunk.choices[0].delta.content or "", end="")
    Full model cardcurated from the model's HuggingFace card — description, training, usage, limitations, reported benchmarks

    Qwen2.5-14B-Instruct is the instruction-tuned 14B model in the Qwen2.5 series from Alibaba Cloud, a family of base and instruction-tuned LLMs ranging from 0.5B to 72B parameters. The card describes it as a causal language model built on the transformer architecture with RoPE, SwiGLU, RMSNorm, and Attention QKV bias. It has 14.7B total parameters (13.1B non-embedding), 48 layers, and grouped-query attention with 40 query heads and 8 key/value heads. The card highlights improvements over prior Qwen releases including more knowledge, stronger coding and mathematics ability, better instruction following, long-text generation over 8K tokens, and improved handling of structured data. It supports a context length of up to 131,072 (128K) tokens, generation up to 8,192 tokens, and multilingual use across more than 29 languages.

    Training data

    The card only classifies the training stage as "Pretraining & Post-training." It does not disclose training datasets, data composition, or the specific fine-tuning methodology used for the instruction-tuned model.

    Intended use
    • +Instruction-tuned conversational assistant used via the transformers chat template (apply_chat_template) with system/user/assistant roles
    • +Coding and mathematics tasks (card cites improved coding and math capabilities)
    • +Long-form text generation exceeding 8K tokens
    • +Structured-data tasks such as understanding tables and generating structured outputs (e.g. JSON)
    • +Long-context workloads up to 128K tokens, with YaRN rope_scaling enabled when inputs exceed 32,768 tokens
    • +Multilingual applications across 29+ languages
    Limitations
    • −Requires transformers >= 4.37.0; older versions raise a KeyError for the qwen2 model type
    • −Long-context (>32,768 tokens) requires manually enabling YaRN rope_scaling in config; the card advises adding it only when long inputs are actually needed
    • −vLLM only supports static YaRN, so the scaling factor stays constant regardless of input length, which the card notes may impact performance on shorter texts
    • −The card does not document safety, bias, or out-of-scope-use guidance beyond the above technical notes
    Running the weights yourself

    The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.

    from transformers import AutoModelForCausalLM, AutoTokenizer
    
    model_name = "Qwen/Qwen2.5-14B-Instruct"
    
    model = AutoModelForCausalLM.from_pretrained(
        model_name,
        torch_dtype="auto",
        device_map="auto"
    )
    tokenizer = AutoTokenizer.from_pretrained(model_name)
    
    prompt = "Give me a short introduction to large language model."
    messages = [
        {"role": "system", "content": "You are Qwen, created by Alibaba Cloud. You are a helpful assistant."},
        {"role": "user", "content": prompt}
    ]
    text = tokenizer.apply_chat_template(
        messages,
        tokenize=False,
        add_generation_prompt=True
    )
    model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
    
    generated_ids = model.generate(
        **model_inputs,
        max_new_tokens=512
    )
    generated_ids = [
        output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
    ]
    
    response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
    Scores

    No measurements published for this version yet.

    Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.

    Our research on this modelwhat we found, and what nobody has measured
    Qwen2.5-14B: a fast, Apache-2.0 base model — read the label before you deploy it

    Alibaba's 14.7B dense checkpoint ships strong self-reported benchmarks and a 128K context, but it is a pretrained base model — not instruction-tuned, and with zero security-specific evaluation. Here is what that means for a practitioner.

    Read the analysis →
    Versionsscores attach to a version; v2 does not inherit v1's numbers
    v12026-07-31—current

    Start with Qwen 2.5 14B

    A confirmed email account includes 30 free messages a month.

    Sign in to start
    Qwen 2.5 14B — General helper · AdversariaLLM