For beginners
Getting started with
local LLMs on a Mac
For newcomers to local LLMs: what Mac you need, how to start with Ollama, the first model to try for your memory, and what this site's numbers mean, based on our measurements.
What a local LLM is
A chat AI like ChatGPT that runs inside your own Mac instead of on an internet service. You download the model (the AI itself) as a file and run it on your Mac.
- Your text stays on your Mac: nothing you type leaves it
- No fees or limits: use it as much as you like, even offline
- The trade-off: not as smart as the big cloud AIs, and how smart and fast it is depends on your Mac's memory and chip
What you need
A Mac with Apple silicon (M1 or later) and a few GB to about 20GB of free disk space for the model file. What matters most is memory: a model that does not fit in memory will not run. As a rule of thumb, a model using up to half your Mac's memory runs comfortably alongside other apps.
Models by memory size: 8GB · 16GB · 24GB · 32GB · 36GB · 48GB or more
First steps (Ollama)
- Install the Mac version of Ollama from ollama.com and start it (free).
- Open Terminal and type the line from the table below. The first time, the model downloads.
- When it finishes, you can chat right there. Type
/byeto quit.
The first model to try, by memory
| Mac memory | First model and the line to type | Memory used | Reply time |
|---|---|---|---|
| 8GB | MiniCPM5-2Bollama run hf.co/openbmb/MiniCPM5-2B-GGUF:Q4_K_M | 2GB | about 1.5 s |
| 16GB | Ornith-1.0-9Bollama run hf.co/ornith-ai/Ornith-1.0-9B-GGUF:Q4_K_M | 6.2GB | about 3.6 s |
| 24GB | Ornith-1.0-9Bollama run hf.co/ornith-ai/Ornith-1.0-9B-GGUF:Q4_K_M | 6.2GB | about 3.6 s |
| 32GB | Ornith-1.0-9Bollama run hf.co/ornith-ai/Ornith-1.0-9B-GGUF:Q4_K_M | 6.2GB | about 3.6 s |
| 36GB | Ornith-1.0-9Bollama run hf.co/ornith-ai/Ornith-1.0-9B-GGUF:Q4_K_M | 6.2GB | about 3.6 s |
| 48GB or more | gemma4:26bollama run gemma4:26b | 18.7GB | about 3 s |
Among the models that fit comfortably, the smartest that replies within 5 seconds (the text page's "if unsure" rule). Reply times are from this site's Mac Studio (M5 Max); slower chips take longer.
Reading this site's numbers
| Number | Meaning |
|---|---|
| Reply time | Seconds until a simple question is fully answered (the middle of 50). Models that think first take longer, as the thinking counts. |
| Knowledge | The share of right answers on Japanese four-choice questions at high-school to university level. Guessing gets 25%. |
| Math | The share of Japanese word problems answered with the right number. |
| Memory used | Memory used while running. Up to half your Mac's memory runs comfortably alongside other apps. |
| Translation score | Out of 100: how many characters match a professional translation. Even a human translation scores below 100; use it to compare models. |
| Margin (±) | With a limited number of questions, the numbers have a range: "80% (±6)" means roughly 74-86%. A gap within the margin means about the same. |
To learn more
- What names like "gemma-4-E4B-it" or "Q4_K_M" mean → how to read a model name
- Words like token or quantization → glossary
- Compare the measured models → text page · translation page