MiniCPM-V
Image + text
Vision helper model for image and text input.
OpenBMB's MiniCPM-V 2.6: a compact multimodal model with strong OCR and multi-image/video understanding for its size.
- ›OCR of a screenshot into text you can act on
- ›Reading multi-page or multi-image evidence
- ›Vision Q&A where OCR accuracy matters
Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AD_API_KEY"],
base_url="https://adversariallm.ai/v1",
)
stream = client.chat.completions.create(
model="minicpm-v",
messages=[{"role": "user", "content": "…"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")MiniCPM-V 2.6 is the latest and most capable model in the MiniCPM-V series, positioned by its authors as a "GPT-4V Level MLLM for Single Image, Multi Image and Video on Your Phone." It is built on SigLip-400M (vision encoder) and Qwen2-7B (language model) for a total of about 8B parameters. It takes images, videos, and text as input and produces text output. Compared to the previous MiniCPM-Llama3-V 2.5, it adds multi-image and video understanding, in-context few-shot learning, and improved OCR. A notable design point is token efficiency: it encodes a 1.8M-pixel image into only 640 visual tokens (the card claims ~75% fewer than most models), which supports on-device / end-side deployment on devices like iPads and phones. It also offers multilingual support.
The card states the model builds on the RLAIF-V and VisCPM techniques and improves over MiniCPM-Llama3-V 2.5. RLAIF-V (reinforcement learning from AI feedback for vision) and the associated RLAIF-V-Dataset are cited in connection with reducing hallucination and improving trustworthy behavior. The card does not enumerate the full pretraining/instruction-tuning corpus beyond these referenced techniques.
- +Single-image understanding and visual question answering
- +Multi-image understanding and reasoning (e.g. comparisons across images)
- +Video understanding (captioning / QA over video frames)
- +OCR and document / scene-text analysis
- +In-context few-shot learning from image-text examples
- +Multilingual use (English, Chinese, German, French, Italian, Korean, and others per the card)
- +On-device / end-side deployment on phones, iPads and similar hardware, enabled by its high token efficiency
- −The card states: as an LMM, the model generates content by learning from a large amount of multimodal corpora, but it cannot comprehend, express personal opinions, or make value judgements.
- −The developers disclaim liability for problems arising from use, including data security issues, risk of public opinion, or any risks and problems arising from misuse, misguidance, illegal use, or misinformation.
- −Hallucination is reduced (the card cites lower rates than GPT-4o/GPT-4V on Object HalBench) but not eliminated.
- −Commercial use requires completing a registration questionnaire; weights are otherwise governed by the MiniCPM Model License.
The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.
import torch
from PIL import Image
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained('openbmb/MiniCPM-V-2_6',
trust_remote_code=True, attn_implementation='sdpa',
torch_dtype=torch.bfloat16)
model = model.eval().cuda()
tokenizer = AutoTokenizer.from_pretrained(
'openbmb/MiniCPM-V-2_6', trust_remote_code=True)
image = Image.open('xx.jpg').convert('RGB')
question = 'What is in the image?'
msgs = [{'role': 'user', 'content': [image, question]}]
res = model.chat(image=None, msgs=msgs, tokenizer=tokenizer)
print(res)Reported by the model’s authors, not our own testing — our scores are in the table above.
| OpenCompass (average, single-image, 8 benchmarks) | 65.2 | vendor-reported, not independently verified |
| OCRBench | claimed state-of-the-art (card says it surpasses GPT-4o, GPT-4V, Gemini 1.5 Pro); no exact score transcribed in the card text | vendor-reported, not independently verified |
| Object HalBench (hallucination) | claimed significantly lower hallucination rate than GPT-4o and GPT-4V; no exact score in card text | vendor-reported, not independently verified |
| Multi-image (Mantis-Eval, BLINK, Mathverse mv, Sciverse mv) | claimed state-of-the-art; numbers shown only in chart images, not transcribed | vendor-reported, not independently verified |
| Video-MME (video understanding) | claimed to outperform GPT-4V, Claude 3.5 Sonnet, and LLaVA-NeXT-Video-34B (with/without subtitles); numbers shown only in charts | vendor-reported, not independently verified |
No measurements published for this version yet.
Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.
The architecture and provenance are clearly documented and the card reports benchmark figures — but the performance claims are vendor-only, with no independent evaluation of this specific version found.
This grades the RECORD, not the model. A model rated E may be excellent — the claim is only that nobody has shown it.
OpenBMB's first-generation MiniCPM-V is a ~3B bilingual vision-language model — a SigLip-400M encoder on a MiniCPM-2.4B LLM. Its benchmarks are self-reported, it has no security tuning, and the famous "GPT-4V on your phone" paper is about a later model, not this one. Here is what that means if you want a local screenshot/OCR helper.
Read the analysis →Start with MiniCPM-V
A confirmed email account includes 30 free messages a month.