# LLaVA-1.5-7B: a well-benchmarked open VLM with zero security evals

*D. Rose · 16 August 2026 · 6 min*

> A CLIP-plus-Vicuna vision-language model with peer-reviewed CVPR 2024 benchmarks — all general-purpose, none security-specific. Here is what it is, how to run it, and exactly where the evidence stops.

LLaVA-1.5-7B is one of the better-documented open vision-language models you can pull today. Its numbers are real, published in a peer-reviewed venue, and reported for the model itself rather than borrowed from a backbone. It is also a general-purpose research model with no security evaluation of any kind. Both of those things are true, and for a security team the second one matters as much as the first.

## What it actually is

LLaVA-v1.5-7B is an auto-regressive multimodal model built by visual instruction tuning ([model card](https://huggingface.co/liuhaotian/llava-v1.5-7b)). The architecture is three parts stitched together ([project repo](https://github.com/haotian-liu/LLaVA)):

- an OpenAI **CLIP ViT-L/14 vision encoder** at 336px resolution,
- a **2-layer MLP projector** that maps vision features into the language space (upgraded from the single linear layer in the original LLaVA), and
- a **Vicuna-v1.5-7B language backbone**.

The "7B" refers to the Vicuna LLM only. The CLIP vision tower (~0.3B) and the MLP projector are additional, and the card does not state a single combined parameter count ([model card](https://huggingface.co/liuhaotian/llava-v1.5-7b)). Worth keeping straight when you size hardware or compare against a text-only 7B.

The provenance chain is clean and fully open: **Llama 2 7B → Vicuna-v1.5-7B → LLaVA-1.5-7B**, with CLIP ViT-L/14 as the vision encoder and the MLP as the connector ([repo](https://github.com/haotian-liu/LLaVA), [model card](https://huggingface.co/liuhaotian/llava-v1.5-7b)). This repo (`liuhaotian/llava-v1.5-7b`) is the original full-precision fp16 PyTorch release — the foundational model itself, not a downstream community quant or fine-tune.

Training used roughly 1.2M public samples ([model card](https://huggingface.co/liuhaotian/llava-v1.5-7b)): 558K BLIP-captioned image-text pairs from LAION/CC/SBU for feature-alignment pretraining, then 158K GPT-generated multimodal instruction samples, ~450K academic-task VQA, and 40K ShareGPT for instruction tuning. It was trained around September 2023; the paper notes the 13B variant trains in about a day on a single 8×A100 node ([Improved Baselines with Visual Instruction Tuning](https://arxiv.org/abs/2310.03744)).

The model is built and maintained by Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee (University of Wisconsin–Madison and Microsoft Research) ([paper](https://arxiv.org/abs/2310.03744)).

## Makers' claims vs. what's verified

This is the section that matters most for this audience, and here the honest answer is short: **the authors make no security claims at all.**

Their stated purpose is deliberately narrow. The model card says "the primary use of LLaVA is research on large multimodal models and chatbots," with primary intended users being "researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence" ([model card](https://huggingface.co/liuhaotian/llava-v1.5-7b)). The paper positions LLaVA-1.5 as a strong, data-efficient open baseline that reached state-of-the-art at the time across 11–12 multimodal benchmarks using only publicly available data ([paper](https://arxiv.org/abs/2310.03744)).

What is verified: the benchmark numbers are published in the peer-reviewed CVPR 2024 paper ([CVPR 2024 proceedings](https://openaccess.thecvf.com/content/CVPR2024/papers/Liu_Improved_Baselines_with_Visual_Instruction_Tuning_CVPR_2024_paper.pdf)) and are reported *for LLaVA-1.5-7B itself* — not silently inherited from Vicuna or Llama 2. That is a real credibility signal, and better than most community VLMs offer. One caveat: these are still the authors' own reported figures, and the Hugging Face card itself publishes no numbers — it only states that 12 benchmarks were used ([model card](https://huggingface.co/liuhaotian/llava-v1.5-7b)). If you need the scores, cite the paper, not the card.

What is **not** verified, because nobody claimed it: any security, malware, threat-intel, or red-team capability. There is no security-specific evaluation of this model, from the authors or anyone else. Any use of LLaVA-1.5-7B in a security workflow is inference-time repurposing of a general research VLM, not a documented capability. Treat it accordingly.

## Benchmarks

Every number below is reported for **the model itself (LLaVA-1.5-7B)** — none is a base-model figure borrowed from Vicuna or Llama 2 — and all trace to the paper ([arXiv:2310.03744v2](https://arxiv.org/html/2310.03744v2)), not the model card. All are general multimodal benchmarks. None is a security task.

| Benchmark | Score | Reported for |
|---|---|---|
| VQAv2 (test-dev) | 78.5 | LLaVA-1.5-7B (the model) |
| GQA | 62.0 | LLaVA-1.5-7B (the model) |
| VizWiz | 50.0 | LLaVA-1.5-7B (the model) |
| ScienceQA-IMG | 66.8 | LLaVA-1.5-7B (the model) |
| TextVQA | 58.2 | LLaVA-1.5-7B (the model) |
| POPE | 85.9 (avg of random/popular/adversarial F1: 87.3 / 86.1 / 84.2) | LLaVA-1.5-7B (the model) |
| MME (perception) | 1510.7 | LLaVA-1.5-7B (the model) |
| MMBench | 64.3 | LLaVA-1.5-7B (the model) |
| MMBench-CN | 58.3 | LLaVA-1.5-7B (the model) |
| SEED-Bench (all) | 66.1 | LLaVA-1.5-7B (the model) |
| LLaVA-Bench (in-the-Wild) | 65.4 | LLaVA-1.5-7B (the model) |
| MM-Vet | 31.1 | LLaVA-1.5-7B (the model) |

Two comparison hazards to note before you drop these into a table next to another model. **POPE** is often quoted as a single 85.9; that is the average across the paper's random/popular/adversarial F1 splits, and third-party papers report POPE F1 anywhere from ~80 to ~86 depending on split and decoding settings — so state the setting when you compare ([arXiv:2310.03744v2](https://arxiv.org/html/2310.03744v2)). **SEED-Bench** here is the "all" figure (66.1); some tables instead report SEED-Bench-Image, which is a different sub-metric ([arXiv:2310.03744v2](https://arxiv.org/html/2310.03744v2)).

## For a practitioner

**What it's genuinely good for.** As a vision helper in a security pipeline, LLaVA-1.5-7B is a capable general describer: screenshots, dashboards, scene content, charts, and diagrams. It will read text in images, but its OCR is modest — TextVQA 58.2 ([arXiv:2310.03744v2](https://arxiv.org/html/2310.03744v2)) — so it is a poor fit as your primary text-extraction stage against dense console output, log screenshots, or fine-print artifacts. Use it to triage and summarize imagery, not to transcribe it verbatim.

**How to run it** ([model card](https://huggingface.co/liuhaotian/llava-v1.5-7b), [repo](https://github.com/haotian-liu/LLaVA)):

- **Weights/format:** original release is fp16 PyTorch weights (~14 GB), served via the official `haotian-liu/LLaVA` repo. Transformers-native conversions exist (`llava-hf`, class `LlavaForConditionalGeneration`), as does vLLM support for easier serving.
- **Hardware:** runs on a single ~16 GB+ GPU in fp16. Community 4-bit/8-bit and GGUF quants exist for llama.cpp/Ollama (via the separate `mmproj` CLIP file) — but those are separate repos, not this one, and no quant-quality numbers are published here.
- **Image input:** one 336×336 image per turn, which produces 576 visual tokens.
- **Context:** 2048-token text window. This is short by current standards and is the practical ceiling — long multi-turn sessions or document-heavy prompts will not fit.
- **Prompt template:** Vicuna-v1.1 style.

**The real caveats.** Beyond the modest OCR and the 2048-token window: there is **no safety or red-team tuning documented** for this model ([model card](https://huggingface.co/liuhaotian/llava-v1.5-7b)). That is a neutral statement of fact, not a claim that it is "uncensored" — it simply means the model ships with whatever behavior the Vicuna/Llama 2 lineage carries and no adversarial hardening or evaluation was performed. If your workflow needs guaranteed refusal behavior or a measured attack-resistance figure, this model does not provide one.

**Licensing.** The weights are under the **Llama 2 Community License Agreement**, inherited through the Vicuna/Llama 2 backbone ([model card](https://huggingface.co/liuhaotian/llava-v1.5-7b)). The LLaVA source code is Apache 2.0, and the project notes you must also comply with underlying licenses including the OpenAI Terms of Use, because GPT-generated data was used in instruction tuning ([repo](https://github.com/haotian-liu/LLaVA)). Read those before any commercial deployment.

## Honest limits

What we do not know, stated plainly:

- **No security evaluation exists** — no malware, phishing, exploit, or threat-intel-imagery benchmark, from the authors or independently ([honest gap, per model card scope](https://huggingface.co/liuhaotian/llava-v1.5-7b)). Repurposing is on you to validate.
- **No combined parameter count** is published; "7B" is the Vicuna backbone only and excludes the ~0.3B CLIP tower and the projector ([model card](https://huggingface.co/liuhaotian/llava-v1.5-7b)).
- **The card publishes no benchmark numbers** — everything in the table above comes from the paper, so a card-only reader would see none of it ([model card](https://huggingface.co/liuhaotian/llava-v1.5-7b), [paper](https://arxiv.org/abs/2310.03744)).
- **Benchmark comparability** needs care: the POPE split and the SEED-Bench sub-metric both change the number ([arXiv:2310.03744v2](https://arxiv.org/html/2310.03744v2)).
- The exact release date of this specific repo is approximate — training dates to ~September 2023 and LLaVA-1.5 was announced in October 2023 ([paper](https://arxiv.org/abs/2310.03744)).

## Bottom line

LLaVA-1.5-7B is a solid, fully-open, peer-reviewed general vision-language model with benchmarks you can actually cite and reproduce. For security work it is a general-purpose vision helper you would repurpose at inference time — good for describing and triaging imagery, weak at OCR, short on context, and carrying zero security-specific validation. If you want a VLM whose credentials are honest, this is one; if you want one whose credentials are *security* credentials, they do not exist for this model.

It sits in the AdversariaLLM measured catalog as [LLaVA-1.5-7B](/catalog/llava-7b), and it is one of the models covered in our [2026 survey of open security-research LLMs](/blog/open-security-research-llms-2026).
