By memory size
Local LLMs for a Mac
with 36GB of memory
On a Mac with 36GB of memory, models up to 18GB run comfortably alongside other apps: 20 of the 30 models measured here. The smartest is muse-glimmer:30b (85% on knowledge, 88% on math, 17.6GB).
Picks by use
| Use | Model | Rule |
|---|---|---|
| If unsure, start here | Ornith-1.0-9B | The smartest that replies within 5 seconds (1 more within the margin) |
| Smartest (if you can wait) | muse-glimmer:30b | The highest share of right answers (2 more within the margin) |
| Replies right away | Mistral-Small-3.2-24B-Instruct-2506 | The smartest that replies within 1 second (3 more within the margin) |
| English → Japanese translation | muse-glimmer:30b | The highest translation score (9 more within the margin) |
Runs comfortably (up to 18GB)
| Model | Memory | Reply time | Knowledge | Math | EN→JA |
|---|---|---|---|---|---|
| muse-glimmer:30b | 17.6GB | about 14.3 s | 85% | 88% | 36 pts |
| Ornith-1.0-9B | 6.2GB | about 3.6 s | 83% | 88% | — |
| gpt-oss:20b | 12.8GB | about 1.9 s | 81% | 86% | 33 pts |
| NVIDIA-Nemotron-Nano-12B-v2 | 8.1GB | about 7.1 s | 79% | 85% | 31 pts |
| gemma-4-E4B-it | 5.4GB | about 4 s | 82% | 79% | 34 pts |
| Ornith-1.5-9B | 6.4GB | about 2.5 s | 78% | 82% | 33 pts |
| Mistral-Small-3.2-24B-Instruct-2506 | 16.3GB | about 0.2 s | 77% | 81% | 34 pts |
| Ministral-3-14B-Instruct-2512 | 9.9GB | about 0.1 s | 72% | 83% | 31 pts |
| Mistral-Small-3.1-24B-Instruct-2503 | 16.3GB | about 0.2 s | 72% | 81% | 34 pts |
| DeepSeek-R1-0528-Qwen3-8B | 6.4GB | about 9.8 s | 73% | 78% | 25 pts |
| MiniCPM5-2B | 2GB | about 1.5 s | 70% | 80% | 23 pts |
| Devstral-Small-2-24B-Instruct-2512 | 16.3GB | about 0.2 s | 66% | 82% | 33 pts |
| LFM2.5-2.6B | 2GB | about 2.2 s | 73% | 68% | 30 pts |
| LFM2.5-8B-A1B | 5.4GB | about 1.5 s | 67% | 71% | 28 pts |
| NVIDIA-Nemotron-3-Nano-4B-BF16 | 3.3GB | about 2.4 s | 63% | 75% | 28 pts |
| GLM-4-9B-0414 | 6.7GB | about 0.1 s | 66% | 64% | 32 pts |
| Qwen3.5-4B | 3.5GB | about 40 s | 56% | 66% | — |
| LFM2-24B-A2B | 14.8GB | about 0.1 s | 54% | 58% | 33 pts |
| LFM2.5-1.2B-JP-202606 | 1GB | about 0.2 s | 57% | 54% | 33 pts |
| Ministral-3-3B-Instruct-2512 | 3.2GB | about 0.1 s | 48% | 42% | 18 pts |
Runs with heavy apps closed
More than half the memory, but up to 23.4GB should run. Close heavy apps other than a browser first.
| Model | Memory | Reply time | Knowledge | Math | EN→JA |
|---|---|---|---|---|---|
| qwen3.6:35b-a3b | 22.3GB | about 8.8 s | 88% | 89% | 29 pts |
| Ornith-1.0-35B | 21.6GB | about 10.1 s | 88% | 88% | 33 pts |
| qwen3.8:27b | 18.2GB | about 6.9 s | 85% | 89% | 31 pts |
| Ornith-1.5-35B-A3B | 21.8GB | about 1.9 s | 82% | 88% | 34 pts |
| DeepSeek-R1-Distill-Qwen-32B | 22.4GB | about 16.8 s | 81% | 86% | 31 pts |
| gemma4:26b | 18.7GB | about 2.9 s | 82% | 85% | 34 pts |
| gemma3:27b | 18.4GB | about 0.4 s | 77% | 84% | 33 pts |
| glm-4.7-flash:latest | 19.5GB | about 6.3 s | 77% | 83% | 32 pts |
| qwen3-coder:30b | 19.4GB | about 0.1 s | 69% | 79% | 31 pts |
| NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 | 22.7GB | about 2 s | 77% | 61% | 31 pts |
About this guide
- Memory was measured on a Mac Studio (M5 Max, 36GB). Memory use is about the same on any Apple silicon Mac, so it carries over.
- Reply times are from this Mac. Slower chips (a MacBook Air, say) reply more slowly.
- "Comfortable" means the model takes at most half the memory; "with heavy apps closed" up to 65%. The picks follow the same rules as the text page, among the models that fit.