Measured model
Measured on a Mac Studio (M5 Max, 36GB)Qwen3-4B
Alibaba · Ollama name hf.co/Qwen/Qwen3-4B-GGUF:Q4_K_M
On a Mac Studio (M5 Max, 36GB), Qwen3-4B answers a simple question in about 3.1 s, gets 72% of the Japanese knowledge questions right, 81% of the math problems, and uses 3.9GB of memory.
Results
| Measure | Result | Among measured models | Note |
|---|---|---|---|
| Reply time | about 3.1 s | 21 of 35 | Until one simple common-sense question is answered (median of 50) |
| Knowledge (high school to university) | 72% | 23 of 35 | 200 questions, ±7 |
| Math (word problems) | 81% | 20 of 35 | 100 questions, ±9 |
| Memory used | 3.9GB | 8 of 35 | |
| EN→JA translation | 29 pts | 22 of 31 | out of 100 (chrF), about 7.2 s a sentence |
| JA→EN translation | 52 pts | 24 of 31 | out of 100 (chrF), about 5.4 s a sentence |
"Among measured models" is the place among the 35 models measured on this site (lower is better for time and memory).
Try it on your Mac
- Install the Mac version of Ollama from ollama.com and start it.
- Type this in Terminal. The first time, the model downloads; then you can chat with it.
ollama run hf.co/Qwen/Qwen3-4B-GGUF:Q4_K_M
It used 3.9GB here. To use it comfortably alongside other apps, a Mac with 8GB of memory or more is a good guide (the model then takes about half the memory).
About the model
- Maker
- Alibaba
- Size
- 4.02B parameters
- Quantization
- Q4_K_M
- Thinks before answering
- Yes
- Longest input at once
- 40,960 tokens
- Released (Hugging Face)
- Apr 27, 2025
- Generation speed
- 123.6 tok/s
- Start-up (loading)
- about 0.5 s
- Measured on
- Oct 7, 2026
- Ollama
- 0.34.3
- Hugging Face
- Qwen/Qwen3-4B
- Series
- Qwen
What the parts of a name mean (size, quantization and so on): how to read a model name.
Models of similar memory
- Qwen3.5-4B (3.5GB of memory, knowledge 81%, reply about 29.7 s)
- NVIDIA-Nemotron-3-Nano-4B-BF16 (3.3GB of memory, knowledge 72%, reply about 1.9 s)
- Ministral-3-3B-Instruct-2512 (3.2GB of memory, knowledge 47%, reply about 0.1 s)
- LFM2.5-8B-A1B (5.4GB of memory, knowledge 71%, reply about 1.7 s)
- gemma-4-E4B-it (5.4GB of memory, knowledge 79%, reply about 2.6 s)
Compare the main models on the text page. The series page lists every model measured.