Measured model
Measured on a Mac Studio (M5 Max, 36GB)MiniCPM5-2B
OpenBMB · Ollama name hf.co/openbmb/MiniCPM5-2B-GGUF:Q4_K_M
On a Mac Studio (M5 Max, 36GB), MiniCPM5-2B answers a simple question in about 1.5 s, gets 70% of the Japanese knowledge questions right, 80% of the math problems, and uses 2GB of memory.
These numbers are from before the method changed on Oct 2, 2026 (see how it was measured).
Results
| Measure | Result | Among measured models | Note |
|---|---|---|---|
| Reply time | about 1.5 s | 11 of 30 | Until one simple common-sense question is answered (median of 50) |
| Knowledge (high school to university) | 70% | 21 of 30 | 200 questions, ±7, 7 cut off (counted wrong) |
| Math (word problems) | 80% | 19 of 30 | 100 questions, ±9, 3 cut off (counted wrong) |
| Memory used | 2GB | 3 of 30 | |
| EN→JA translation | 23 pts | 27 of 28 | out of 100 (chrF), about 9 s a sentence |
| JA→EN translation | 48 pts | 27 of 28 | out of 100 (chrF), about 8.3 s a sentence |
"Among measured models" is the place among the 30 models measured on this site (lower is better for time and memory).
Try it on your Mac
- Install the Mac version of Ollama from ollama.com and start it.
- Type this in Terminal. The first time, the model downloads; then you can chat with it.
ollama run hf.co/openbmb/MiniCPM5-2B-GGUF:Q4_K_M
It used 2GB here. To use it comfortably alongside other apps, a Mac with 8GB of memory or more is a good guide (the model then takes about half the memory).
About the model
- Maker
- OpenBMB
- Size
- 2.52B parameters
- Quantization
- Q4_K_M
- Thinks before answering
- Yes
- Longest input at once
- 131,072 tokens
- Released (Hugging Face)
- Sep 6, 2026
- Generation speed
- 186.2 tok/s
- Start-up (loading)
- about 0.8 s
- Measured on
- Sep 26, 2026
- Ollama
- 0.34.3
- Hugging Face
- openbmb/MiniCPM5-2B
- Series
- MiniCPM
Models of similar memory
- LFM2.5-2.6B (2GB of memory, knowledge 73%, reply about 2.2 s)
- LFM2.5-1.2B-JP-202606 (1GB of memory, knowledge 57%, reply about 0.2 s)
- Ministral-3-3B-Instruct-2512 (3.2GB of memory, knowledge 48%, reply about 0.1 s)
- NVIDIA-Nemotron-3-Nano-4B-BF16 (3.3GB of memory, knowledge 63%, reply about 2.4 s)
- Qwen3.5-4B (3.5GB of memory, knowledge 56%, reply about 40 s)
Compare the main models on the text page. The series page lists every model measured.