It delivers outstanding overall performance among open-source base models with fewer than 4 billion parameters; equipped with capabilities such as tool use, code generation, and long-context reasoning, it demonstrates the initial potential to serve as a general-purpose on-device agent.
Highlights
53.9 AA Comprehensive Score
Artificial Analysis — Top Performance Among Sub-4B Open Models:On the AA index, MiniCPM5-2B's comprehensive score reaches 23, ranking first among open-source models under 4B parameters — surpassing Gemma 4 12B, Qwen3.5 9B, and other 3B–4B models.
Early Signs of On-Device General Agent:Achieves a breakthrough score of 20 on agentic benchmarks — 10× ahead of same-tier models — demonstrating early high-density on-device general Agent capability.
Data Governance for Smarter Intelligence:Built on UltraData's L0–L4 tiered data governance, with four newly open-sourced datasets: UltraData-Code, UltraData-SFT-Agent-2609, UltraData-RL-2609, and UltraX.
Technical highlights
UltraData-Driven Full-Stage Training:MiniCPM5-2B is a complete realization of the UltraData tiered data management system, covering Base Training, Mid-Training, and Post-Training. Training data comes from the co-released Ultra-FineWeb, Ultra-FineWeb-L3, UltraData-Math, and UltraData-SFT-2605, spanning general knowledge, math, code, reasoning, and agentic tasks.
RL + OPD: Significant Gains in Reasoning & Agent:RL + OPD is the key post-training stage. The RL phase uses a critic-based algorithm, achieving average gains of ↑10.96 on reasoning and general tasks, and ↑6.96 on Agent tasks. The OPD phase merges 16 RL expert models — reusing their training prompts as distillation data without constructing additional corpora.
Benchmarks
Overall:In overall evaluation (average score), MiniCPM5-2B reaches SOTA among similarly-sized open-source models (average 53.9), surpassing every larger model in the comparison (highest 51.1).
Reasoning:Across code and math reasoning benchmarks, MiniCPM5-2B leads across the board among similarly-sized models.
Knowledge:On general knowledge benchmarks, MiniCPM5-2B leads similarly-sized models on MMLU-Pro, MMLU-Redux and more.
Instruction Following:On instruction-following benchmarks, MiniCPM5-2B leads similarly-sized models on IFBench and Multi-IF.
Long Context:On long-context benchmarks, MiniCPM5-2B leads on AA-LCR, NoLiMa and more.
Agent:Across agentic benchmarks — tool use, coding agents, search agents and general agents — MiniCPM5-2B takes a commanding lead.
* Data from real model evaluations; † Scores taken from Artificial Analysis official figures; all others are internally reproduced results