Have 30-200 Employees? Make $50k-$500k selling your data for AI Training.

Learn More

MiniCPM5-2B Ranks First Among Open-Weight Models Under 4B

Naveed Wali Khan
Naveed Wali Khan
Published in

The AI briefing for Developers

Stay up to date with AI tools, model releases, and developer workflows that matter.

Weekly. Free. One click to leave.

Share this article

MiniCPM5-2B Ranks First Among Open-Weight Models Under 4B
SitePoint Premium
Stay Relevant and Grow Your Career in Tech
  • Premium Results
  • Publish articles on SitePoint
  • Daily curated jobs
  • Learning Paths
  • Discounts to dev tools
Start Free Trial

7 Day Free Trial. Cancel Anytime.

Artificial Analysis runs a benchmark called GDPval-AA v2. It scores models on real-world work tasks and anchors the scale to a human baseline of 1,000.

MiniCPM5-2B scores 831 and finishes first. Ling 3.0 Tiny follows at 718, then Granite 4.2 8B at 648, Gemma 4 12B at 593, and Qwen3.5 9B at 590.

With 2.6 billion parameters, MiniCPM5-2B is the smallest model in that field and it sits at the top of it.

On the composite Intelligence Index that GDPval-AA v2 feeds into, it places third. The two models ahead of it are substantially larger: Ling 3.0 Tiny carries roughly three times the parameters, and Gemma 4 12B is several times larger again. Reading those two results side by side is the most useful thing in this release, and it says something about how to read benchmark scores that applies well beyond one model.

Capability Density, Properly Stated

In December 2024, researchers at Tsinghua and ModelBest proposed a second dimension to measure models along. Scaling laws describe how performance improves as you add parameters and data. Alongside that, they argued, it is worth tracking capability per parameter and watching how that ratio moves over time.

They called it capability density. Take a model, work out the parameter count a reference model would need to match it, and divide.

Estimated capability density of open-source base LLMs over time
Fig: estimated capability density of open-source base LLMs. Figure covers releases through September 2024. Source: Xiao et al., Densing Law of LLMs.

Plotted over time, the frontier rises exponentially. Llama-1, from February 2023, sits below 0.1. By September 2024, Gemma-2-9B and MiniCPM-3-4B are close to 2. Fitted across 51 open-source pre-trained foundation models and five benchmarks, the maximum capability density envelope corresponds to a doubling time of roughly 3.5 months.

Two things about that fit need stating carefully. It describes the frontier, not any individual model, so it does not mean a given capability can be reproduced with half the parameters every three and a half months. And it was measured on base models, with instruction tuning, retrieval and inference-time scaling deliberately excluded to avoid confounders. MiniCPM5-2B is post-trained, so it is not a data point on that curve.

Where It Lands on v4.2

Artificial Analysis Intelligence Index v4.2 scatter plot against parameter count
Fig: Artificial Analysis Intelligence Index v4.2. Dotted line is the Pareto frontier; shaded area is Artificial Analysis's most attractive quadrant. Data provisional pending completion of independent evaluation.

This plots parameters on a log axis against the Intelligence Index. The dotted line is the Pareto frontier: the set of models where nothing smaller scores higher.

MiniCPM5-2B sits on that frontier and is the smallest model on it in this field. The segment from there across to Ling 3.0 Tiny is almost flat: in this comparison Ling 3.0 Tiny has roughly three times the total parameters and scores one index point higher.

Read across the chart at that height and you find six models, from 2.6B up to Gemma 4 12B, clustered within two index points of each other. Read below it and Qwen3 30B A3B 2507 sits six points lower. In this selected comparison, models of similar size show substantial variation, suggesting that parameter count alone is an increasingly weak predictor of benchmark performance.

Artificial Analysis Intelligence Index v4.2 bar chart across ten evaluations
Fig: Artificial Analysis Intelligence Index v4.2, incorporating 10 evaluations. Hatched bars are estimates pending independent evaluation. Figures remain subject to Artificial Analysis's official testing.

On the index itself MiniCPM5-2B scores 15, behind Ling 3.0 Tiny and Gemma 4 12B at 16 and level with Qwen3.5 9B. That makes it first among open-weight models under 4B parameters. All figures here are Intelligence Index v4.2, which incorporates ten evaluations including AA-Briefcase and GDPval-AA v2. Scores are specific to an index version and are not comparable across versions.

This is not confined to models you can run locally. Z.ai's GLM-5.3-Flash has 320 billion total parameters and 18 billion active, and it beats GLM-5.2 across their published benchmarks at roughly a tenth of the price, including 63.4 against 46.2 on DeepSWE v1.1. Same argument, two orders of magnitude up the parameter axis.

The Composite Index Is the Wrong Lens

A composite score is an average of things you may not want averaged. Intelligence Index v4.2 contains agentic benchmarks, coding benchmarks, and knowledge benchmarks including Humanity's Last Exam and AA-Omniscience. At this scale, broad factual recall is generally harder to match against substantially larger models, and MiniCPM5-2B does give ground on several knowledge-heavy evaluations.

Split the index into its parts and the ordering changes.

Elo rating for real-world work tasks anchored to a human baseline
Fig: Elo rating for real-world work tasks, anchored to a human baseline of 1,000. One of ten evaluations in Artificial Analysis Intelligence Index v4.2.

On GDPval-AA v2, MiniCPM5-2B is first at 831, 113 Elo clear of Ling 3.0 Tiny and 183 clear of Granite 4.2 8B. The error bars are visible on the chart and the margin is much larger than any of them.

Agentic knowledge work benchmark results
Fig: agentic knowledge work benchmark from Artificial Analysis, aggregating rubric pass rate, analytical quality Elo and presentation Elo. One of ten evaluations in Intelligence Index v4.2.

On AA-Briefcase, an agentic knowledge work benchmark, it comes second at 438 behind Ling 3.0 Tiny at 485, still ahead of Granite 4.2 8B at 331 and Granite 4.2 3B at 134.

So: third on the composite, first and second on the two benchmarks that measure getting things done. The index is not wrong, it is measuring a blend, and the composite score masks both the model's strong agentic results and its weaker performance on knowledge-heavy tasks. Both readings come from the same Artificial Analysis data.

What the Reasoning Costs

Weighted average output tokens per task split into answer and reasoning
Fig: weighted average output tokens to run one task in Artificial Analysis Intelligence Index v4.2, split into answer and reasoning.

Reasoning models pay for their reasoning in output tokens, and that cost is uneven.

Ling 3.0 Tiny scores 16 and spends 56k output tokens per task. MiniCPM5-2B scores 15 and spends 19k. Granite 4.2 8B scores 14 on 33k. Granite 4.2 3B scores 11 on the same 19k as MiniCPM5-2B.

That last pairing is the clearest one. Identical token budget, identical task set, four index points apart.

These are the numbers as published rather than a derived efficiency metric, and output tokens are not the same thing as inference cost, since they say nothing about active parameters, input processing, cache behaviour or wall-clock time on real hardware. But if you are running a loop that reasons repeatedly on a device you own, the difference between 19k and 56k tokens per task is something you will feel.

What You Are Actually Running

Architecturally there is not much to report, which is the point. MiniCPM5-2B is a dense causal language model on the standard LlamaForCausalLM layout, 42 layers deep, with grouped-query attention across 16 query heads and 2 key-value heads. Context length is 131,072 tokens, and the licence is Apache-2.0. The name comes from the 1.98 billion non-embedding parameters; counting embeddings the total is about 2.6 billion, which is the figure charts plot it at.

Memory follows from that. The weights need roughly 5 GB at BF16, 2.5 GB at INT8, and 1.26 GB at INT4, and that last number is usually the one that decides whether a model fits the hardware you already own.

OpenBMB is upfront that there are no reproducible throughput benchmarks for this model on smartphone or laptop CPUs yet, so treat any speed figure as something to measure on your own hardware.

The Profile, and Where It Gives Ground

Radar chart of full benchmark results from the official model card
Fig: full results as published on the official model card. Blue marks the best result in each row across all models; black marks the best among 2B-class models. Radar axes are scaled per dimension, so spoke length shows relative position rather than absolute score.

OpenBMB's own evaluation table puts MiniCPM5-2B at an average of 53.9 across 34 benchmarks, against Qwen3.5-4B at 51.1 and granite-4.2-3B at 42.7. The table marks its own provenance: rows carrying a dagger are Artificial Analysis's official numbers, and the rest are reproduced internally.

The 2B-class comparison is the easy one. The interesting comparison is Qwen3.5-4B, and against that model MiniCPM5-2B is not uniformly better. It is specialised.

It leads on all five coding rows, including LCB-Pro Medium at 17.5 where the whole 2B class scores zero. It leads both AIME years, all three tool-use rows, all three search-agent rows, three of four general-agent rows, and NoLiMa by 68.1 to 43.5.

It gives ground on every general knowledge row: MMLU-Pro 70.8 to 78.0, GPQA-Diamond 70.2 to 77.1, SuperGPQA 40.8 to 52.8. It loses MATH-500, all three instruction-following rows, three of four long-context rows, and SWE-bench Pro. On instruction following it is still the strongest 2B-class model in the table, but granite-4.2-3B takes both IFBench and IFEval outright. On Terminal-Bench v2.1 it scores 8.6 against 25.8, the sharpest single gap in the table. Long-horizon terminal work remains a weakness for MiniCPM5-2B in this evaluation.

The knowledge gap is the one most likely to matter in practice, and it is partly addressable. Retrieval and search can mitigate part of the factual-coverage gap, although the model must still select, interpret, and combine the retrieved evidence correctly. Calibration, meaning whether a model knows when to stop, is a separate problem that still requires its own training, evaluation and safeguards. Both are worth planning for before you put a local model inside an autonomous loop, because errors there compound across turns.

How It Was Trained, and Why You Can Check

Post-training runs in three stages. First, 400B tokens of deep-thinking supervised fine-tuning. Then reinforcement learning, training specialised teachers for maths, code, agentic tasks and writing. Then On-Policy Distillation, which merges 16 expert models produced by that RL stage, five of them agentic, back into a single released model, using full-vocabulary reverse KL divergence between student and teacher logits as the advantage estimate.

Score gains from RL and On-Policy Distillation over the SFT baseline
Fig: score gains from RL + OPD, showing the SFT baseline and the increment added by the RL and distillation stages.

The gains are not evenly distributed, which is the interesting part. GPQA-Diamond goes from 48.59 to 70.2. LCB-Pro Easy from 45.36 to 68.04. AIME 2025 from 66.46 to 86.46. SWE-bench Verified from 29 to 46.4. Averaged out, OpenBMB reports 10.96 points on reasoning and general capability and 6.96 on agentic capability.

The part worth paying attention to is that the machinery is public. The day after the model, OpenBMB released Meshy, the RL framework it was trained with.

Diagram of one Meshy training cycle across rollout, inference and training services
Fig: one Meshy training cycle. Rollout, inference and training run as independent services over a shared data plane, with a gate signal pacing generation against training.

Meshy is service-based: it splits inference, rollout and training into separate services communicating over a unified data plane, rather than running them as one process. Rollout requests generations over HTTP from an SGLang inference service, tags each sample with the weight version that produced it, and writes it to a transfer queue. Training pulls rows once every column is present, steps, saves a checkpoint, and emits a gate signal that paces how far ahead rollout is allowed to run.

That design is what lets the same framework support on-policy, bounded off-policy and fully asynchronous RL by changing how tight the gate is. It is built on SGLang and TorchTitan, uses recipe-driven SPMD placement, and ships ready-made recipes for MiniCPM5-1B and 2B including 128K context-parallel training. The credit-assignment method is documented separately as JustRL II, a critic-based approach aimed at token-level credit assignment across long reasoning chains.

Meshy is at github.com/OpenBMB/Meshy.

Four datasets ship with the model. UltraX is a function-calling approach to data refinement, turning fine-grained edits into executable operations, and its open preview contains around 100B tokens across five refined web datasets. Alongside it are UltraData-Code with tiered code data, UltraData-SFT-Agent-2609 with 500K agent training samples, and UltraData-RL-2609 with more than 80K reinforcement learning samples. The base, mid-training and SFT-only checkpoints are published alongside the final one.

Taken together, that is the data, the framework, the intermediate checkpoints and the final weights. Enough to test which stage produced which number rather than taking the recipe on trust, which very few labs make possible.

Running It

MiniCPM5-2B uses the standard LlamaForCausalLM architecture and is supported by mainstream inference frameworks without custom model code, though tokenizer, chat template and quantization support still vary by runtime.

pip install "vllm>=0.21"
vllm serve openbmb/MiniCPM5-2B --port 8000

SGLang 0.5.16 or later is the recommended backend for tool calling, since the model emits XML-style tool calls and SGLang's minicpm5 parser converts them to OpenAI-compatible tool_calls directly. Transformers needs 5.6 or later. On NVIDIA GPUs, vLLM and SGLang are the standard paths. GGUF builds cover llama.cpp, Ollama and LM Studio, there is an MLX build for Apple Silicon, and ArcLight handles CPU inference. Fine-tuning works through TRL, LLaMA-Factory, ms-swift and unsloth. Recommended sampling is temperature 1.0 with top_p 0.95.

Weights and cookbooks are at huggingface.co/openbmb/MiniCPM5-2B and github.com/OpenBMB/MiniCPM under Apache-2.0.

If you evaluate it, the knowledge benchmarks will tell you roughly what the table already says. Give it tools and a retrieval index instead and watch what it does with them, because that is the axis it was built for.

What This Means If You Build Things

The useful conclusion is not that a 2.6B model scored 15 on an index.

It is that "how big does the model need to be" has stopped being a good proxy for "will this work". A model that shows strong reasoning and tool-use benchmark performance at this size, and occupies 1.26 GB at INT4, is not a cheaper version of a cloud API. It is a different component, and the architectures worth designing are the ones that treat it as one: local reasoning and tool use, remote knowledge, and nothing leaving the device that does not need to.

The limits are real and easy to state. Factual coverage is thin at this size. Its Terminal-Bench v2.1 score of 8.6 indicates that long-horizon terminal work remains a significant weakness. Artificial Analysis has not finished independent evaluation, and the numbers may move.

Over the historical period the Densing Law paper covers, the fitted frontier trend corresponds to a parameter doubling time of approximately 3.5 months, though it remains inconclusive whether this rate will persist. What is already clear is that the range of things you can run on hardware you own has widened faster than most roadmaps assumed, and it is worth checking whether yours still assumes otherwise.

Naveed Wali KhanNaveed Wali Khan

I write about AI and software development from the angle that usually gets skipped: what a spec sheet or benchmark actually implies once you try to build with it. Find me on LinkedIn.

© 2000 – 2026 SitePoint Pty. Ltd.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.