Everything measured on one Mac Studio, the mothership

On a Mac Studio (M5 Max, 36GB),how well
does text AI run?

Memory
36GB
Chip
M5 Max
CPU
18 cores
GPU
32 cores

The mothership: Mac Studio (model MHL64J/A), macOS 27.0. Every number on this site was measured on this one Mac. Whether an AI runs at all comes down mostly to how much memory you have.

AI like ChatGPT can run entirely on your own Mac (a local LLM). Your text never leaves the machine, and there is no charge for what you use. But models differ a lot in how fast and how smart they are. On this page, popular models are tried on the mothership Mac Studio under the same conditions, and the results are compared in plain terms.

Picks by use

Chosen from the results below by fixed rules. Models within the margin are said to be level.

Side by side

Bold marks the best value in each column. Click a model name to jump to its full results.

Model Reply time Knowledge Math Memory In short
qwen3.6:35b-a3bAlibabaabout 8.8 s88%89%22.3GB
qwen3.8:27bAlibabaabout 6.9 s85%89%18.2GB
muse-glimmer:30bMetaabout 14.3 s85%88%17.6GB
gemma4:26bGoogleabout 2.9 s82%85%18.7GB
gpt-oss:20bOpenAIabout 1.9 s81%86%12.8GB
glm-4.7-flash:latestZ.aiabout 6.3 s77%83%19.5GB
gemma3:27bGoogleabout 0.4 s74%83%18.4GB
MiniCPM5-2BOpenBMBabout 1.5 s70%80%2GB
qwen3-coder:30bAlibabaabout 0.1 s68%82%19.4GB

Results by model

To try it on your own Mac

  1. Install Ollama. Download the Mac version from the official site, ollama.com, move it to your Applications folder and open it.
  2. Pick a model. From the results above, choose one that fits in your Mac's memory. A model that fits in about half your memory stays comfortable even with other apps open.
  3. Start talking. Type the following in Terminal. The first time, the model downloads (10GB or more); once it finishes you can chat. ollama run gpt-oss:20b

How it was measured

  • Smarts: every model answers the same questions from public Japanese question sets, and the share of correct answers is shown. The knowledge questions are 200 four-choice questions from 53 high-school to university subjects; the math set is 100 word problems. There is a cap on how long a model may think. A question it doesn't answer within the cap, or whose answer started repeating itself and was stopped, is "cut off": it counts as wrong, and the number is shown under the score. A question whose reply Ollama couldn't read three times in a row also counts as wrong.
  • Reply time: the time to finish answering one easy common-sense question (JCommonsenseQA, CC BY-SA 4.0; the median of 50). For models that think before answering, the thinking time is included.
  • Memory: the memory used once the model is loaded. The length of text read at once is set to about 8,000 tokens (around 10,000 Japanese characters) for every model.
  • Everything is measured on the same Mac, each model with the settings its maker recommends (temperature and the like); a model with no recommended settings is measured with randomness turned off. The random seed is fixed. Results will differ on other Macs or with other settings (since October 2, 2026; before that every model was measured with randomness turned off, but models that think before answering sometimes kept repeating themselves without end).
  • The mothership measures around the clock, so anything measured while the operator is using the same Mac may come out a little slower.
  • Public question sets may have been used to train the models, so scores can come out higher than their real ability.
  • Measuring again: the same model under the same settings gives nearly the same result, so a model is measured again only when a new version comes out, when the method changes, or when Ollama or macOS is updated (then only the speed is measured again; the scores stay).
  • Models on this page: up to 12. When the page is full and a new model joins, one model, such as a previous generation, moves to the series page (its results stay; the models holding an agent's role and the gpt-oss baseline stay here).
  • Release date: the day the original model appeared on Hugging Face (when its repository was created). Repositories are sometimes created privately before an announcement, so it can be a little earlier than the announced date.
  • Translation results are on the translation page, and the words used are explained in About.