By memory size
Local LLMs for a Mac
with 8GB of memory
On a Mac with 8GB of memory, models up to 4GB run comfortably alongside other apps: 6 of the 30 models measured here. The smartest is MiniCPM5-2B (70% on knowledge, 80% on math, 2GB).
Picks by use
| Use | Model | Rule |
|---|---|---|
| If unsure, start here | MiniCPM5-2B | The smartest that replies within 5 seconds (1 more within the margin) |
| Smartest (if you can wait) | MiniCPM5-2B | The highest share of right answers (1 more within the margin) |
| Replies right away | LFM2.5-1.2B-JP-202606 | The smartest that replies within 1 second |
| English → Japanese translation | LFM2.5-1.2B-JP-202606 | The highest translation score (1 more within the margin) |
Runs comfortably (up to 4GB)
| Model | Memory | Reply time | Knowledge | Math | EN→JA |
|---|---|---|---|---|---|
| MiniCPM5-2B | 2GB | about 1.5 s | 70% | 80% | 23 pts |
| LFM2.5-2.6B | 2GB | about 2.2 s | 73% | 68% | 30 pts |
| NVIDIA-Nemotron-3-Nano-4B-BF16 | 3.3GB | about 2.4 s | 63% | 75% | 28 pts |
| Qwen3.5-4B | 3.5GB | about 40 s | 56% | 66% | — |
| LFM2.5-1.2B-JP-202606 | 1GB | about 0.2 s | 57% | 54% | 33 pts |
| Ministral-3-3B-Instruct-2512 | 3.2GB | about 0.1 s | 48% | 42% | 18 pts |
About this guide
- Memory was measured on a Mac Studio (M5 Max, 36GB). Memory use is about the same on any Apple silicon Mac, so it carries over.
- Reply times are from this Mac. Slower chips (a MacBook Air, say) reply more slowly.
- "Comfortable" means the model takes at most half the memory; "with heavy apps closed" up to 65%. The picks follow the same rules as the text page, among the models that fit.