By memory size
Local LLMs for a Mac
with 16GB of memory
On a Mac with 16GB of memory, models up to 8GB run comfortably alongside other apps: 12 of the 30 models measured here. The smartest is Ornith-1.0-9B (83% on knowledge, 88% on math, 6.2GB).
Picks by use
| Use | Model | Rule |
|---|---|---|
| If unsure, start here | Ornith-1.0-9B | The smartest that replies within 5 seconds |
| Smartest (if you can wait) | Ornith-1.0-9B | The highest share of right answers |
| Replies right away | GLM-4-9B-0414 | The smartest that replies within 1 second |
| English → Japanese translation | gemma-4-E4B-it | The highest translation score (3 more within the margin) |
Runs comfortably (up to 8GB)
| Model | Memory | Reply time | Knowledge | Math | EN→JA |
|---|---|---|---|---|---|
| Ornith-1.0-9B | 6.2GB | about 3.6 s | 83% | 88% | — |
| gemma-4-E4B-it | 5.4GB | about 4 s | 82% | 79% | 34 pts |
| Ornith-1.5-9B | 6.4GB | about 2.5 s | 78% | 82% | 33 pts |
| DeepSeek-R1-0528-Qwen3-8B | 6.4GB | about 9.8 s | 73% | 78% | 25 pts |
| MiniCPM5-2B | 2GB | about 1.5 s | 70% | 80% | 23 pts |
| LFM2.5-2.6B | 2GB | about 2.2 s | 73% | 68% | 30 pts |
| LFM2.5-8B-A1B | 5.4GB | about 1.5 s | 67% | 71% | 28 pts |
| NVIDIA-Nemotron-3-Nano-4B-BF16 | 3.3GB | about 2.4 s | 63% | 75% | 28 pts |
| GLM-4-9B-0414 | 6.7GB | about 0.1 s | 66% | 64% | 32 pts |
| Qwen3.5-4B | 3.5GB | about 40 s | 56% | 66% | — |
| LFM2.5-1.2B-JP-202606 | 1GB | about 0.2 s | 57% | 54% | 33 pts |
| Ministral-3-3B-Instruct-2512 | 3.2GB | about 0.1 s | 48% | 42% | 18 pts |
Runs with heavy apps closed
More than half the memory, but up to 10.4GB should run. Close heavy apps other than a browser first.
| Model | Memory | Reply time | Knowledge | Math | EN→JA |
|---|---|---|---|---|---|
| NVIDIA-Nemotron-Nano-12B-v2 | 8.1GB | about 7.1 s | 79% | 85% | 31 pts |
| Ministral-3-14B-Instruct-2512 | 9.9GB | about 0.1 s | 72% | 83% | 31 pts |
About this guide
- Memory was measured on a Mac Studio (M5 Max, 36GB). Memory use is about the same on any Apple silicon Mac, so it carries over.
- Reply times are from this Mac. Slower chips (a MacBook Air, say) reply more slowly.
- "Comfortable" means the model takes at most half the memory; "with heavy apps closed" up to 65%. The picks follow the same rules as the text page, among the models that fit.