MiniCPM-V

Image + text

Vision helper model for image and text input.

Context8,192
Curationcommunity
AvailabilityAvailable
Price / Mtok$0.15 in · $0.60 out
Quantization—
LicenseMiniCPM Model License — read before use
Good for
    Model cardfrom the published weights — parameters, architecture, training, intended use
    Parameters8.1B
    ArchitectureMiniCPM-V 2.6 (MiniCPMV, vision)
    Modalityvision
    Training

    OpenBMB's MiniCPM-V 2.6: a compact multimodal model with strong OCR and multi-image/video understanding for its size.

    Intended use
    • ›OCR of a screenshot into text you can act on
    • ›Reading multi-page or multi-image evidence
    • ›Vision Q&A where OCR accuracy matters
    Call it from your own codethe API is OpenAI-compatible — an existing client needs one line changed

    Your key comes from /keys. Every request is metered and audited against your account, and the model id is the slug in this page’s address.

    import os
    from openai import OpenAI
    
    client = OpenAI(
        api_key=os.environ["AD_API_KEY"],
        base_url="https://adversariallm.ai/v1",
    )
    
    stream = client.chat.completions.create(
        model="minicpm-v",
        messages=[{"role": "user", "content": "…"}],
        stream=True,
    )
    for chunk in stream:
        print(chunk.choices[0].delta.content or "", end="")
    Full model cardcurated from the model's HuggingFace card — description, training, usage, limitations, reported benchmarks

    MiniCPM-V 2.6 is the latest and most capable model in the MiniCPM-V series, positioned by its authors as a "GPT-4V Level MLLM for Single Image, Multi Image and Video on Your Phone." It is built on SigLip-400M (vision encoder) and Qwen2-7B (language model) for a total of about 8B parameters. It takes images, videos, and text as input and produces text output. Compared to the previous MiniCPM-Llama3-V 2.5, it adds multi-image and video understanding, in-context few-shot learning, and improved OCR. A notable design point is token efficiency: it encodes a 1.8M-pixel image into only 640 visual tokens (the card claims ~75% fewer than most models), which supports on-device / end-side deployment on devices like iPads and phones. It also offers multilingual support.

    Training data

    The card states the model builds on the RLAIF-V and VisCPM techniques and improves over MiniCPM-Llama3-V 2.5. RLAIF-V (reinforcement learning from AI feedback for vision) and the associated RLAIF-V-Dataset are cited in connection with reducing hallucination and improving trustworthy behavior. The card does not enumerate the full pretraining/instruction-tuning corpus beyond these referenced techniques.

    Intended use
    • +Single-image understanding and visual question answering
    • +Multi-image understanding and reasoning (e.g. comparisons across images)
    • +Video understanding (captioning / QA over video frames)
    • +OCR and document / scene-text analysis
    • +In-context few-shot learning from image-text examples
    • +Multilingual use (English, Chinese, German, French, Italian, Korean, and others per the card)
    • +On-device / end-side deployment on phones, iPads and similar hardware, enabled by its high token efficiency
    Limitations
    • −The card states: as an LMM, the model generates content by learning from a large amount of multimodal corpora, but it cannot comprehend, express personal opinions, or make value judgements.
    • −The developers disclaim liability for problems arising from use, including data security issues, risk of public opinion, or any risks and problems arising from misuse, misguidance, illegal use, or misinformation.
    • −Hallucination is reduced (the card cites lower rates than GPT-4o/GPT-4V on Object HalBench) but not eliminated.
    • −Commercial use requires completing a registration questionnaire; weights are otherwise governed by the MiniCPM Model License.
    Running the weights yourself

    The maker’s own snippet, from the model card — it downloads the weights and runs them on your hardware. Kept here because reproducing a result independently is the point, not because you need it to use the model.

    import torch
    from PIL import Image
    from transformers import AutoModel, AutoTokenizer
    
    model = AutoModel.from_pretrained('openbmb/MiniCPM-V-2_6',
        trust_remote_code=True, attn_implementation='sdpa',
        torch_dtype=torch.bfloat16)
    model = model.eval().cuda()
    tokenizer = AutoTokenizer.from_pretrained(
        'openbmb/MiniCPM-V-2_6', trust_remote_code=True)
    
    image = Image.open('xx.jpg').convert('RGB')
    question = 'What is in the image?'
    msgs = [{'role': 'user', 'content': [image, question]}]
    
    res = model.chat(image=None, msgs=msgs, tokenizer=tokenizer)
    print(res)
    Benchmarks — as reported on the card

    Reported by the model’s authors, not our own testing — our scores are in the table above.

    OpenCompass (average, single-image, 8 benchmarks)65.2vendor-reported, not independently verified
    OCRBenchclaimed state-of-the-art (card says it surpasses GPT-4o, GPT-4V, Gemini 1.5 Pro); no exact score transcribed in the card textvendor-reported, not independently verified
    Object HalBench (hallucination)claimed significantly lower hallucination rate than GPT-4o and GPT-4V; no exact score in card textvendor-reported, not independently verified
    Multi-image (Mantis-Eval, BLINK, Mathverse mv, Sciverse mv)claimed state-of-the-art; numbers shown only in chart images, not transcribedvendor-reported, not independently verified
    Video-MME (video understanding)claimed to outperform GPT-4V, Claude 3.5 Sonnet, and LLaVA-NeXT-Video-34B (with/without subtitles); numbers shown only in chartsvendor-reported, not independently verified
    Scores

    No measurements published for this version yet.

    Baseline is the strongest general-purpose model we could run on the same suite, same setup, same day. The control row tells you what the other rows are worth.

    Evidencewhat is actually known about this artifact

    The architecture and provenance are clearly documented and the card reports benchmark figures — but the performance claims are vendor-only, with no independent evaluation of this specific version found.

    This grades the RECORD, not the model. A model rated E may be excellent — the claim is only that nobody has shown it.

    Our research on this modelwhat we found, and what nobody has measured
    MiniCPM-V (first-gen): a 3B bilingual vision helper with self-reported numbers and no security evals

    OpenBMB's first-generation MiniCPM-V is a ~3B bilingual vision-language model — a SigLip-400M encoder on a MiniCPM-2.4B LLM. Its benchmarks are self-reported, it has no security tuning, and the famous "GPT-4V on your phone" paper is about a later model, not this one. Here is what that means if you want a local screenshot/OCR helper.

    Read the analysis →
    Versionsscores attach to a version; v2 does not inherit v1's numbers
    v12026-07-31—current

    Start with MiniCPM-V

    A confirmed email account includes 30 free messages a month.

    Sign in to start
    MiniCPM-V — Vision helper · AdversariaLLM