Measured model
Measured on a Mac Studio (M5 Max, 36GB)glm-4.7-flash:latest
Z.ai
On a Mac Studio (M5 Max, 36GB), glm-4.7-flash:latest answers a simple question in about 6.3 s, gets 77% of the Japanese knowledge questions right, 83% of the math problems, and uses 19.5GB of memory.
These numbers are from before the method changed on Oct 2, 2026 (see how it was measured).
Results
| Measure | Result | Among measured models | Note |
|---|---|---|---|
| Reply time | about 6.3 s | 22 of 30 | Until one simple common-sense question is answered (median of 50) |
| Knowledge (high school to university) | 77% | 13 of 30 | 200 questions, ±6, 21 cut off (counted wrong) |
| Math (word problems) | 83% | 11 of 30 | 100 questions, ±9, 7 cut off (counted wrong) |
| Memory used | 19.5GB | 25 of 30 | |
| EN→JA translation | 32 pts | 14 of 28 | out of 100 (chrF), about 23.3 s a sentence |
| JA→EN translation | 53 pts | 21 of 28 | out of 100 (chrF), about 19.9 s a sentence |
"Among measured models" is the place among the 30 models measured on this site (lower is better for time and memory).
Try it on your Mac
- Install the Mac version of Ollama from ollama.com and start it.
- Type this in Terminal. The first time, the model downloads; then you can chat with it.
ollama run glm-4.7-flash:latest
It used 19.5GB here. To use it comfortably alongside other apps, a Mac with 48GB of memory or more is a good guide (the model then takes about half the memory).
About the model
- Maker
- Z.ai
- Size
- 29.9B parameters
- Quantization
- Q4_K_M
- Thinks before answering
- Yes
- Longest input at once
- 202,752 tokens
- Released (Hugging Face)
- Jan 19, 2026
- Generation speed
- 80.5 tok/s
- Start-up (loading)
- about 9.1 s
- Measured on
- Sep 26, 2026
- Ollama
- 0.34.3
- Hugging Face
- zai-org/GLM-4.7-Flash
- Series
- GLM
Models of similar memory
- qwen3-coder:30b (19.4GB of memory, knowledge 68%, reply about 0.1 s)
- gemma4:26b (18.7GB of memory, knowledge 82%, reply about 2.9 s)
- gemma3:27b (18.4GB of memory, knowledge 74%, reply about 0.4 s)
- qwen3.8:27b (18.2GB of memory, knowledge 85%, reply about 6.9 s)
- muse-glimmer:30b (17.6GB of memory, knowledge 85%, reply about 14.3 s)
Compare the main models on the text page. The series page lists every model measured.