# MiniCPM-V (first-gen): a 3B bilingual vision helper with self-reported numbers and no security evals

*D. Rose · 16 August 2026 · 5 min*

> OpenBMB's first-generation MiniCPM-V is a ~3B bilingual vision-language model — a SigLip-400M encoder on a MiniCPM-2.4B LLM. Its benchmarks are self-reported, it has no security tuning, and the famous "GPT-4V on your phone" paper is about a later model, not this one. Here is what that means if you want a local screenshot/OCR helper.

MiniCPM-V is a small, bilingual vision-language model. It is useful, it is honest about being general-purpose, and almost everything you will read about "MiniCPM-V" online is about a *later* model with the same name. This post keeps those apart.

## What it actually is

MiniCPM-V is OpenBMB's first-generation vision-language model, released 1 February 2024 and also labelled OmniLMM-3B. It is a ~3B-parameter multimodal model that bolts a vision encoder onto a small language model. The model card describes it as "built based on SigLip-400M and MiniCPM-2.4B, connected by a perceiver resampler," and notes that image inputs are compressed "into 64 tokens via a perceiver resampler" ([model card README](https://huggingface.co/openbmb/MiniCPM-V/blob/main/README.md)).

The provenance chain is a composite, not a single lineage:

- **Vision encoder:** Google's SigLip-400M.
- **Language backbone:** OpenBMB's MiniCPM-2B (~2.4B parameters excluding embeddings, from ModelBest Inc. and TsinghuaNLP).
- **Connector:** a perceiver resampler that compresses vision tokens down to 64.

Both halves and the connector are stated on the [model card README](https://huggingface.co/openbmb/MiniCPM-V/blob/main/README.md). The creator is OpenBMB, a collaboration of ModelBest Inc. and TsinghuaNLP (the Tsinghua University NLP lab) ([GitHub](https://github.com/openbmb/MiniCPM-V)).

Weights are full-precision BF16. The model is bilingual (English and Chinese), and the card states it runs on GPU, CPU, and mobile (Android/Harmony) ([model card](https://huggingface.co/openbmb/MiniCPM-V)). This repository is not itself a GGUF or quantized build; int4/GGUF variants exist chiefly for the later versions.

One point does most of the work in this post: **this is the *first* generation.** It is distinct from MiniCPM-Llama3-V 2.5 (May 2024), 2.6 (Aug 2024), 4.0/4.5 (2025), and 4.6 (2026), which use different, larger backbones and high-resolution any-aspect-ratio encoding ([GitHub](https://github.com/openbmb/MiniCPM-V)). Capabilities from those releases do not transfer to this one.

## Makers' claims vs. what is verified

The card makes one headline performance claim, and it is self-reported. OpenBMB states MiniCPM-V "achieves the best overall performance among multimodal models of the same scale, surpassing existing multimodal large models built on Phi-2 and achieving performance comparable to or even better than 9.6B Qwen-VL-Chat on some tasks" ([model card](https://huggingface.co/openbmb/MiniCPM-V)). That is the author's positioning, not an independent finding.

The declared Hugging Face task is Visual Question Answering ([model card](https://huggingface.co/openbmb/MiniCPM-V)). There is no security, malware, threat-intel, or adversarial framing anywhere on the card. The stated purpose is general bilingual image understanding on constrained hardware.

**Watch the paper trail.** Two arXiv papers get associated with the string "MiniCPM-V," and neither is an independent benchmark of *this* model:

- The arXiv link on this model's own card, [2308.12038](https://arxiv.org/abs/2308.12038), is actually the **VisCPM** paper ("Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages") — a related OpenBMB multimodal work, not a MiniCPM-V benchmark paper.
- The widely cited ["MiniCPM-V: A GPT-4V Level MLLM on Your Phone" (2408.01800)](https://arxiv.org/abs/2408.01800), from August 2024, primarily evaluates the **later** MiniCPM-Llama3-V 2.5. Its GPT-4V / Gemini / Claude-3-beating OpenCompass results belong to that model, not to this 3B first-gen build. Do not attribute them here.

If you cite a "MiniCPM-V beats GPT-4V" number for this repo, you are citing the wrong model.

## Benchmarks

Every number below is **self-reported by OpenBMB on the model card**, and every one is for **this first-generation model itself** — not a base component, and not a later version. No independent third-party reproduction was found for this specific release. Read them as vendor-published figures.

| Benchmark | Score | Model or base? | Source |
|---|---|---|---|
| MME | 1452 | The model (MiniCPM-V first-gen) — self-reported | [card README](https://huggingface.co/openbmb/MiniCPM-V/blob/main/README.md) |
| MMBench dev (English) | 67.9 | The model (MiniCPM-V first-gen) — self-reported | [card README](https://huggingface.co/openbmb/MiniCPM-V/blob/main/README.md) |
| MMBench dev (Chinese) | 65.3 | The model (MiniCPM-V first-gen) — self-reported | [card README](https://huggingface.co/openbmb/MiniCPM-V/blob/main/README.md) |
| MMMU val | 37.2 | The model (MiniCPM-V first-gen) — self-reported | [card README](https://huggingface.co/openbmb/MiniCPM-V/blob/main/README.md) |
| CMMMU val | 32.1 | The model (MiniCPM-V first-gen) — self-reported | [card README](https://huggingface.co/openbmb/MiniCPM-V/blob/main/README.md) |

There are no security-relevant benchmarks — none published, none found. There are also no independent numbers of any kind for this version.

## For a practitioner

**How to run it.** Load via Hugging Face Transformers with `trust_remote_code=True`. The card lists Python 3.10+, torch 2.1.2, transformers 4.36.0, timm 0.9.10, and Pillow 10.1.0 ([model card](https://huggingface.co/openbmb/MiniCPM-V)). Because vision inputs are compressed to 64 tokens, the multimodal footprint stays small enough for CPU and mobile inference ([card README](https://huggingface.co/openbmb/MiniCPM-V/blob/main/README.md)) — which is the real reason to reach for it over a larger VLM.

**What you would genuinely use it for.** A local, bilingual (EN/ZH) vision helper: describing a screenshot, captioning an image inside a triage pipeline, or answering questions about a picture on an air-gapped or resource-constrained box where a bigger model will not fit. That is the honest scope. It is a vision helper, not a security driver — it carries no security-specific fine-tuning and no evaluation on security-relevant tasks ([card README](https://huggingface.co/openbmb/MiniCPM-V/blob/main/README.md)).

**Real caveats.**

- **It is the oldest MiniCPM-V.** If you need stronger OCR, high-resolution any-aspect-ratio input (up to ~1.8M px), or document/PDF parsing, those belong to the later versions (2.5 / 2.6 / 4.x), not this 3B first-gen. Its OCR and high-resolution behavior are comparatively limited ([card README](https://huggingface.co/openbmb/MiniCPM-V/blob/main/README.md)).
- **The license is split and use-restricted.** The code is Apache-2.0, but the weights are governed by the "MiniCPM Model License": free for academic research, with commercial use permitted only after questionnaire registration ([card README](https://huggingface.co/openbmb/MiniCPM-V/blob/main/README.md)). This is not an OSI-approved open-weights license — clear it before any commercial deployment.
- **No dual-use story, and that is worth stating plainly.** This is an ordinary general-purpose model. There is no refusal-removal, no abliteration, and no offensive-security framing anywhere in the research on it. Its risk profile is unremarkable; the caveat here is capability, not safety.

## Honest limits: what is not known

- **Context length and image resolution are not pinned down.** The card we read does not state the first-gen model's exact context/token length or input image resolution. A web search hinted at ~448px, but that could not be confirmed from a primary source, so we do not assert it.
- **The self-reported numbers are unreproduced.** No independent evaluation of this specific version was found.
- **No security-domain evaluation exists** for malware, phishing, exploit, or threat-intel imagery — none claimed by the maker, none found elsewhere.
- **Quantization status is unconfirmed** from the card. This repo ships BF16; whether a first-gen GGUF/int4 variant exists was not verified (those repos exist mainly for later versions).

Overall evidence strength for this model is **moderate**: the architecture and provenance are clearly documented, but the performance claims are vendor-only and the security-relevant record is empty.

## In the catalog

MiniCPM-V is tracked in the AdversariaLLM catalog at [/catalog/minicpm-v](/catalog/minicpm-v), where models are listed by measured, cited evidence rather than maker claims. For how it fits the wider landscape of open security-research models, see [Open Security-Research LLMs, 2026](/blog/open-security-research-llms-2026).
