Measured model
Measured on a Mac Studio (M5 Max, 36GB)gemma-4-12B-it
Google · Ollama name hf.co/unsloth/gemma-4-12b-it-GGUF:Q4_K_M
On a Mac Studio (M5 Max, 36GB), gemma-4-12B-it answers a simple question in about 7.5 s, gets 85% of the Japanese knowledge questions right, 87% of the math problems, and uses 8.1GB of memory.
Results
| Measure | Result | Among measured models | Note |
|---|---|---|---|
| Reply time | about 7.5 s | 23 of 31 | Until one simple common-sense question is answered (median of 50) |
| Knowledge (high school to university) | 85% | 5 of 31 | 200 questions, ±6, 1 cut off (counted wrong) |
| Math (word problems) | 87% | 8 of 31 | 100 questions, ±8, 2 cut off (counted wrong) |
| Memory used | 8.1GB | 13 of 31 | |
| EN→JA translation | 35 pts | 2 of 29 | out of 100 (chrF), about 18 s a sentence |
| JA→EN translation | 58 pts | 2 of 29 | out of 100 (chrF), about 17.3 s a sentence |
"Among measured models" is the place among the 31 models measured on this site (lower is better for time and memory).
Try it on your Mac
- Install the Mac version of Ollama from ollama.com and start it.
- Type this in Terminal. The first time, the model downloads; then you can chat with it.
ollama run hf.co/unsloth/gemma-4-12b-it-GGUF:Q4_K_M
It used 8.1GB here. To use it comfortably alongside other apps, a Mac with 24GB of memory or more is a good guide (the model then takes about half the memory).
About the model
- Maker
- Size
- 11.9B parameters
- Quantization
- Q4_K_M
- Thinks before answering
- No
- Longest input at once
- 262,144 tokens
- Released (Hugging Face)
- May 23, 2026
- Generation speed
- 48.7 tok/s
- Start-up (loading)
- about 1.1 s
- Measured on
- Oct 3, 2026
- Ollama
- 0.34.3
- Hugging Face
- google/gemma-4-12B-it
- Series
- Gemma
What the parts of a name mean (size, quantization and so on): how to read a model name.
Models of similar memory
- NVIDIA-Nemotron-Nano-12B-v2 (8.1GB of memory, knowledge 79%, reply about 7.1 s)
- GLM-4-9B-0414 (6.7GB of memory, knowledge 66%, reply about 0.1 s)
- DeepSeek-R1-0528-Qwen3-8B (6.4GB of memory, knowledge 73%, reply about 9.8 s)
- Ornith-1.5-9B (6.4GB of memory, knowledge 78%, reply about 2.5 s)
- Ministral-3-14B-Instruct-2512 (9.9GB of memory, knowledge 72%, reply about 0.1 s)
Compare the main models on the text page. The series page lists every model measured.