LLaVA 7B
Image + text
Vision helper model for image and text input.
LLaVA v1.5-7B: a Vicuna/Llama base joined to a CLIP vision encoder and instruction-tuned on image-text pairs, so it can read a screenshot as well as text.
- ›Reading a screenshot of an advisory, console or config
- ›Describing a diagram or attack graph
- ›General image-plus-question tasks
Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AD_API_KEY"],
base_url="https://adversariallm.ai/v1",
)
stream = client.chat.completions.create(
model="llava-7b",
messages=[{"role": "user", "content": "…"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")LLaVA is an open-source chatbot trained by fine-tuning LLaMA/Vicuna on GPT-generated multimodal instruction-following data. It is an auto-regressive language model based on the transformer architecture, targeting image-text-to-text (vision-language) tasks. This checkpoint, LLaVA-v1.5-7B, was trained in September 2023. The card states its primary use is research on large multimodal models and chatbots, aimed at researchers and hobbyists in computer vision, NLP, and machine learning. It is distributed under the Llama 2 Community License.
Per the card's "Training dataset" section: 558K filtered image-text pairs from LAION/CC/SBU, captioned by BLIP; 158K GPT-generated multimodal instruction-following data; 450K academic-task-oriented VQA data mixture; and 40K ShareGPT data. Method: fine-tuning of a LLaMA/Vicuna base language model on GPT-generated multimodal instruction-following data.
- +Primary intended use: research on large multimodal models and chatbots
- +Primary intended users: researchers and hobbyists in computer vision, natural language processing, machine learning, and artificial intelligence
- +Questions/comments directed to the LLaVA GitHub issues tracker (github.com/haotian-liu/LLaVA/issues); more info at llava-vl.github.io
No measurements published for this version yet.
Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.
A CLIP-plus-Vicuna vision-language model with peer-reviewed CVPR 2024 benchmarks — all general-purpose, none security-specific. Here is what it is, how to run it, and exactly where the evidence stops.
Read the analysis →Start with LLaVA 7B
A confirmed email account includes 30 free messages a month.