Measured model
Measured on a Mac Studio (M5 Max, 36GB)Ornith-1.0-35B
ornith-ai · Ollama name hf.co/ornith-ai/Ornith-1.0-35B-GGUF:Q4_K_M
On a Mac Studio (M5 Max, 36GB), Ornith-1.0-35B answers a simple question in about 10.1 s, gets 88% of the Japanese knowledge questions right, 88% of the math problems, and uses 21.6GB of memory.
These numbers are from before the method changed on Oct 2, 2026 (see how it was measured).
Results
| Measure | Result | Among measured models | Note |
|---|---|---|---|
| Reply time | about 10.1 s | 27 of 30 | Until one simple common-sense question is answered (median of 50) |
| Knowledge (high school to university) | 88% | 2 of 30 | 200 questions, ±5 |
| Math (word problems) | 88% | 3 of 30 | 100 questions, ±8 |
| Memory used | 21.6GB | 26 of 30 | |
| EN→JA translation | 33 pts | 10 of 28 | out of 100 (chrF), about 19.2 s a sentence |
| JA→EN translation | 57 pts | 8 of 28 | out of 100 (chrF), about 18.2 s a sentence |
"Among measured models" is the place among the 30 models measured on this site (lower is better for time and memory).
Try it on your Mac
- Install the Mac version of Ollama from ollama.com and start it.
- Type this in Terminal. The first time, the model downloads; then you can chat with it.
ollama run hf.co/ornith-ai/Ornith-1.0-35B-GGUF:Q4_K_M
It used 21.6GB here. To use it comfortably alongside other apps, a Mac with 48GB of memory or more is a good guide (the model then takes about half the memory).
About the model
- Maker
- ornith-ai
- Size
- 34.7B parameters
- Quantization
- Q4_K_M
- Thinks before answering
- Yes
- Longest input at once
- 262,144 tokens
- Released (Hugging Face)
- Jun 21, 2026
- Generation speed
- 98.9 tok/s
- Start-up (loading)
- about 1.3 s
- Measured on
- Sep 27, 2026
- Ollama
- 0.34.3
- Hugging Face
- ornith-ai/Ornith-1.0-35B
- Series
- Ornith
Models of similar memory
- Ornith-1.5-35B-A3B (21.8GB of memory, knowledge 82%, reply about 1.9 s)
- qwen3.6:35b-a3b (22.3GB of memory, knowledge 88%, reply about 8.8 s)
- DeepSeek-R1-Distill-Qwen-32B (22.4GB of memory, knowledge 81%, reply about 16.8 s)
- NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 (22.7GB of memory, knowledge 77%, reply about 2 s)
- glm-4.7-flash:latest (19.5GB of memory, knowledge 77%, reply about 6.3 s)
Compare the main models on the text page. The series page lists every model measured.