MiniCPM-V (first-gen): a 3B bilingual vision helper with self-reported numbers and no security evals
D. Rose · 16 August 2026 · 5 min
OpenBMB's first-generation MiniCPM-V is a ~3B bilingual vision-language model — a SigLip-400M encoder on a MiniCPM-2.4B LLM. Its benchmarks are self-reported, it has no security tuning, and the famous "GPT-4V on your phone" paper is about a later model, not this one. Here is what that means if you want a local screenshot/OCR helper.
MiniCPM-V is a small, bilingual vision-language model. It is useful, it is honest about being general-purpose, and almost everything you will read about "MiniCPM-V" online is about a later model with the same name. This post keeps those apart.
What it actually is
MiniCPM-V is OpenBMB's first-generation vision-language model, released 1 February 2024 and also labelled OmniLMM-3B. It is a ~3B-parameter multimodal model that bolts a vision encoder onto a small language model. The model card describes it as "built based on SigLip-400M and MiniCPM-2.4B, connected by a perceiver resampler," and notes that image inputs are compressed "into 64 tokens via a perceiver resampler" (model card README).
The provenance chain is a composite, not a single lineage:
- Vision encoder: Google's SigLip-400M.
- Language backbone: OpenBMB's MiniCPM-2B (~2.4B parameters excluding embeddings, from ModelBest Inc. and TsinghuaNLP).
- Connector: a perceiver resampler that compresses vision tokens down to 64.
Both halves and the connector are stated on the model card README. The creator is OpenBMB, a collaboration of ModelBest Inc. and TsinghuaNLP (the Tsinghua University NLP lab) (GitHub).
Weights are full-precision BF16. The model is bilingual (English and Chinese), and the card states it runs on GPU, CPU, and mobile (Android/Harmony) (model card). This repository is not itself a GGUF or quantized build; int4/GGUF variants exist chiefly for the later versions.
One point does most of the work in this post: this is the first generation. It is distinct from MiniCPM-Llama3-V 2.5 (May 2024), 2.6 (Aug 2024), 4.0/4.5 (2025), and 4.6 (2026), which use different, larger backbones and high-resolution any-aspect-ratio encoding (GitHub). Capabilities from those releases do not transfer to this one.
Makers' claims vs. what is verified
The card makes one headline performance claim, and it is self-reported. OpenBMB states MiniCPM-V "achieves the best overall performance among multimodal models of the same scale, surpassing existing multimodal large models built on Phi-2 and achieving performance comparable to or even better than 9.6B Qwen-VL-Chat on some tasks" (model card). That is the author's positioning, not an independent finding.
The declared Hugging Face task is Visual Question Answering (model card). There is no security, malware, threat-intel, or adversarial framing anywhere on the card. The stated purpose is general bilingual image understanding on constrained hardware.
Watch the paper trail. Two arXiv papers get associated with the string "MiniCPM-V," and neither is an independent benchmark of this model:
- The arXiv link on this model's own card, 2308.12038, is actually the VisCPM paper ("Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages") — a related OpenBMB multimodal work, not a MiniCPM-V benchmark paper.
- The widely cited "MiniCPM-V: A GPT-4V Level MLLM on Your Phone" (2408.01800), from August 2024, primarily evaluates the later MiniCPM-Llama3-V 2.5. Its GPT-4V / Gemini / Claude-3-beating OpenCompass results belong to that model, not to this 3B first-gen build. Do not attribute them here.
If you cite a "MiniCPM-V beats GPT-4V" number for this repo, you are citing the wrong model.
Benchmarks
Every number below is self-reported by OpenBMB on the model card, and every one is for this first-generation model itself — not a base component, and not a later version. No independent third-party reproduction was found for this specific release. Read them as vendor-published figures.
| Benchmark | Score | Model or base? | Source |
|---|---|---|---|
| MME | 1452 | The model (MiniCPM-V first-gen) — self-reported | card README |
| MMBench dev (English) | 67.9 | The model (MiniCPM-V first-gen) — self-reported | card README |
| MMBench dev (Chinese) | 65.3 | The model (MiniCPM-V first-gen) — self-reported | card README |
| MMMU val | 37.2 | The model (MiniCPM-V first-gen) — self-reported | card README |
| CMMMU val | 32.1 | The model (MiniCPM-V first-gen) — self-reported | card README |
There are no security-relevant benchmarks — none published, none found. There are also no independent numbers of any kind for this version.
For a practitioner
How to run it. Load via Hugging Face Transformers with trust_remote_code=True. The card lists Python 3.10+, torch 2.1.2, transformers 4.36.0, timm 0.9.10, and Pillow 10.1.0 (model card). Because vision inputs are compressed to 64 tokens, the multimodal footprint stays small enough for CPU and mobile inference (card README) — which is the real reason to reach for it over a larger VLM.
What you would genuinely use it for. A local, bilingual (EN/ZH) vision helper: describing a screenshot, captioning an image inside a triage pipeline, or answering questions about a picture on an air-gapped or resource-constrained box where a bigger model will not fit. That is the honest scope. It is a vision helper, not a security driver — it carries no security-specific fine-tuning and no evaluation on security-relevant tasks (card README).
Real caveats.
- It is the oldest MiniCPM-V. If you need stronger OCR, high-resolution any-aspect-ratio input (up to ~1.8M px), or document/PDF parsing, those belong to the later versions (2.5 / 2.6 / 4.x), not this 3B first-gen. Its OCR and high-resolution behavior are comparatively limited (card README).
- The license is split and use-restricted. The code is Apache-2.0, but the weights are governed by the "MiniCPM Model License": free for academic research, with commercial use permitted only after questionnaire registration (card README). This is not an OSI-approved open-weights license — clear it before any commercial deployment.
- No dual-use story, and that is worth stating plainly. This is an ordinary general-purpose model. There is no refusal-removal, no abliteration, and no offensive-security framing anywhere in the research on it. Its risk profile is unremarkable; the caveat here is capability, not safety.
Honest limits: what is not known
- Context length and image resolution are not pinned down. The card we read does not state the first-gen model's exact context/token length or input image resolution. A web search hinted at ~448px, but that could not be confirmed from a primary source, so we do not assert it.
- The self-reported numbers are unreproduced. No independent evaluation of this specific version was found.
- No security-domain evaluation exists for malware, phishing, exploit, or threat-intel imagery — none claimed by the maker, none found elsewhere.
- Quantization status is unconfirmed from the card. This repo ships BF16; whether a first-gen GGUF/int4 variant exists was not verified (those repos exist mainly for later versions).
Overall evidence strength for this model is moderate: the architecture and provenance are clearly documented, but the performance claims are vendor-only and the security-relevant record is empty.
In the catalog
MiniCPM-V is tracked in the AdversariaLLM catalog at /catalog/minicpm-v, where models are listed by measured, cited evidence rather than maker claims. For how it fits the wider landscape of open security-research models, see Open Security-Research LLMs, 2026.