Measured model
Measured on a Mac Studio (M5 Max, 36GB)GLM-4-9B-0414
Z.ai · Ollama name hf.co/unsloth/GLM-4-9B-0414-GGUF:Q4_K_M
On a Mac Studio (M5 Max, 36GB), GLM-4-9B-0414 answers a simple question in about 0.1 s, gets 66% of the Japanese knowledge questions right, 64% of the math problems, and uses 6.7GB of memory.
These numbers are from before the method changed on Oct 2, 2026 (see how it was measured).
Results
| Measure | Result | Among measured models | Note |
|---|---|---|---|
| Reply time | about 0.1 s | 4 of 30 | Until one simple common-sense question is answered (median of 50) |
| Knowledge (high school to university) | 66% | 24 of 30 | 200 questions, ±7 |
| Math (word problems) | 64% | 26 of 30 | 100 questions, ±10 |
| Memory used | 6.7GB | 12 of 30 | |
| EN→JA translation | 32 pts | 15 of 28 | out of 100 (chrF), about 0.7 s a sentence |
| JA→EN translation | 56 pts | 12 of 28 | out of 100 (chrF), about 0.5 s a sentence |
"Among measured models" is the place among the 30 models measured on this site (lower is better for time and memory).
Try it on your Mac
- Install the Mac version of Ollama from ollama.com and start it.
- Type this in Terminal. The first time, the model downloads; then you can chat with it.
ollama run hf.co/unsloth/GLM-4-9B-0414-GGUF:Q4_K_M
It used 6.7GB here. To use it comfortably alongside other apps, a Mac with 16GB of memory or more is a good guide (the model then takes about half the memory).
About the model
- Maker
- Z.ai
- Size
- 9.4B parameters
- Quantization
- Q4_K_M
- Thinks before answering
- No
- Longest input at once
- 32,768 tokens
- Released (Hugging Face)
- Apr 7, 2025
- Generation speed
- 61.7 tok/s
- Start-up (loading)
- about 0.8 s
- Measured on
- Sep 30, 2026
- Ollama
- 0.34.3
- Hugging Face
- zai-org/GLM-4-9B-0414
- Series
- GLM
Models of similar memory
- DeepSeek-R1-0528-Qwen3-8B (6.4GB of memory, knowledge 73%, reply about 9.8 s)
- Ornith-1.5-9B (6.4GB of memory, knowledge 78%, reply about 2.5 s)
- Ornith-1.0-9B (6.2GB of memory, knowledge 83%, reply about 3.6 s)
- gemma-4-E4B-it (5.4GB of memory, knowledge 82%, reply about 4 s)
- LFM2.5-8B-A1B (5.4GB of memory, knowledge 67%, reply about 1.5 s)
Compare the main models on the text page. The series page lists every model measured.