LLaVA 7B

Image + text

Vision helper model for image and text input.

Context8,192
Curationcommunity
AvailabilityAvailable
Price / Mtok$0.15 in · $0.60 out
BaseVicuna-7B (Llama) + CLIP ViT-L
Quantization—
LicenseLlama 2 Community License
Good for
    Model cardfrom the published weights — parameters, architecture, training, intended use
    Parameters7B
    ArchitectureLLaVA 1.5 (LlavaLlamaForCausalLM, vision)
    Modalityvision
    Training

    LLaVA v1.5-7B: a Vicuna/Llama base joined to a CLIP vision encoder and instruction-tuned on image-text pairs, so it can read a screenshot as well as text.

    Intended use
    • ›Reading a screenshot of an advisory, console or config
    • ›Describing a diagram or attack graph
    • ›General image-plus-question tasks
    Call it from your own codethe API is OpenAI-compatible — an existing client needs one line changed

    Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.

    import os
    from openai import OpenAI
    
    client = OpenAI(
        api_key=os.environ["AD_API_KEY"],
        base_url="https://adversariallm.ai/v1",
    )
    
    stream = client.chat.completions.create(
        model="llava-7b",
        messages=[{"role": "user", "content": "…"}],
        stream=True,
    )
    for chunk in stream:
        print(chunk.choices[0].delta.content or "", end="")
    Full model cardcurated from the model's HuggingFace card — description, training, usage, limitations, reported benchmarks

    LLaVA is an open-source chatbot trained by fine-tuning LLaMA/Vicuna on GPT-generated multimodal instruction-following data. It is an auto-regressive language model based on the transformer architecture, targeting image-text-to-text (vision-language) tasks. This checkpoint, LLaVA-v1.5-7B, was trained in September 2023. The card states its primary use is research on large multimodal models and chatbots, aimed at researchers and hobbyists in computer vision, NLP, and machine learning. It is distributed under the Llama 2 Community License.

    Training data

    Per the card's "Training dataset" section: 558K filtered image-text pairs from LAION/CC/SBU, captioned by BLIP; 158K GPT-generated multimodal instruction-following data; 450K academic-task-oriented VQA data mixture; and 40K ShareGPT data. Method: fine-tuning of a LLaMA/Vicuna base language model on GPT-generated multimodal instruction-following data.

    Intended use
    • +Primary intended use: research on large multimodal models and chatbots
    • +Primary intended users: researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence
    • +Questions/comments directed to the LLaVA GitHub issues tracker (github.com/haotian-liu/LLaVA/issues); more info at llava-vl.github.io
    Scores

    No measurements published for this version yet.

    Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.

    Our research on this modelwhat we found, and what nobody has measured
    LLaVA-1.5-7B: a well-benchmarked open VLM with zero security evals

    A CLIP-plus-Vicuna vision-language model with peer-reviewed CVPR 2024 benchmarks — all general-purpose, none security-specific. Here is what it is, how to run it, and exactly where the evidence stops.

    Read the analysis →
    Versionsscores attach to a version; v2 does not inherit v1's numbers
    v12026-07-31—current

    Start with LLaVA 7B

    A confirmed email account includes 30 free messages a month.

    Sign in to start
    LLaVA 7B — Vision helper · AdversariaLLM