Everything measured on one Mac Studio, the mothership

Local LLMs comparedby series and
generation

A series is a family of models one developer keeps releasing, like Qwen or Gemma. This page lists every generation and size of each series since 2025, measures the ones that run on the mothership Mac Studio (36GB of memory), and shows how much each generation changed.

Words on this page

Series
A family of models one developer keeps releasing (Qwen, Gemma and so on). The list is not fixed: a family that becomes widely used joins it automatically, by fixed rules.
Generation
A version within a series (Qwen3.5, then Qwen3.6, and so on). One generation often comes in several sizes.
Size band
Three bands by the number of parameters: small up to about 5 billion (light even on a laptop), medium up to about 15 billion, and large above that.

All series

Most used first. Select one to jump to its results.

  • Qwen by Alibaba: 4 of 31 models measured newest measured: Qwen3.8 27B, knowledge questions 85%
  • Gemma by Google: 3 of 12 models measured newest measured: gemma-4 26B-A4B, knowledge questions 82%
  • Ornith by ornith-ai: 4 of 6 models measured newest measured: Ornith-1.5 35B-A3B, knowledge questions 82%
  • DeepSeek by DeepSeek: 2 of 20 models measured newest measured: DeepSeek-R1-0528-Qwen3 8B, knowledge questions 73%
  • MiniMax by MiniMax: 0 of 10 models measured
  • Nemotron by NVIDIA: 3 of 8 models measured newest measured: NVIDIA-Nemotron-3 30B-A3B, knowledge questions 77%
  • LFM by Liquid AI: 4 of 13 models measured newest measured: LFM2.5-JP-202606 1.2B, knowledge questions 57%
  • GLM by Z.ai: 2 of 13 models measured newest measured: GLM-4.7 31.2B, knowledge questions 77%
  • Mistral by Mistral: 5 of 16 models measured newest measured: Devstral-2-2512 24B, knowledge questions 66%
  • Llama by Meta: 0 of 2 models measured
  • Muse by Meta: 1 of 1 models measured newest measured: Muse-Glimmer 30B, knowledge questions 85%
  • gpt-oss by OpenAI: 1 of 2 models measured newest measured: gpt-oss 20B, knowledge questions 81%
  • Granite by IBM: 0 of 18 models measured
  • Kimi by Moonshot AI: 0 of 8 models measured
  • MiniCPM by OpenBMB: 1 of 8 models measured newest measured: MiniCPM5 2B, knowledge questions 70%
  • Phi by Microsoft: 0 of 5 models measured
  • SmolLM by Hugging Face: 0 of 6 models measured
  • OLMo by Ai2: 0 of 22 models measured

Generation by generation (knowledge questions)

Results by series

How models are chosen and measured

  • Models listed: every chat model each series' developer has published on Hugging Face since 2025, collected automatically every day. Models only for images, audio or a narrow task, and the same weights in another format, are left out.
  • How widely used: each series' share of the last 30 days' downloads among the 1,000 most downloaded files in GGUF, the format used to run models on a Mac. Downloads from Ollama are not counted (Ollama forbids automated access).
  • Models measured: for series with a share of 0.5% or more, one model per generation and size band: the largest that runs on this Mac. A model on the Text AI page stands for its generation and band. Models whose 4-bit file is over about 65% of the Mac's memory (about 23GB) are not measured.
  • Older generations: this site does not try to measure every version; it keeps looking for models worth running on this Mac now. So in each series and size band, only the newest generation and the most downloaded one (people sometimes still prefer an older version) are measured. Other older generations are marked "Not measured: a newer generation exists". Lines made for a different use, such as coding, long reasoning or Japanese, are counted separately. Results already measured stay after a newer generation comes out.
  • Order: the Text AI page's models come first, then newer generations before older ones. A generation whose measured models answered the questions within 3 hours moves to the front, and one whose model took longer or ran into a time limit goes last. This only brings quick results out sooner; it does not affect the scores. A measured model is removed from the Mac; only its results are kept.
  • Method: the same as the Text AI page (200 knowledge questions, 100 math questions, reply time, memory, English→Japanese translation). Results not yet redone since the method changed are marked "Earlier method".
  • Timed out: each model has a time limit (8 hours for the questions, 2 hours for the translation). A model that runs past it would do the same the next day, so it is not tried again and is marked "Timed out". When the method changes, it is tried once more (on October 2, 2026 the method changed to stop answers that repeat themselves, so the timed-out models are being measured again).
  • Unreadable replies: sometimes a model's reply is malformed and Ollama returns an error instead of the reply. Only that question is then asked again, up to twice. If it still cannot be read, the model's measurement counts as failed and is tried once more the next day.
  • Changes between generations: every model answers the same 200 questions, so two generations are compared question by question. The questions only one of them got right are counted, and "up" or "down" is written only when that split is wider than chance would give once in 20 times; otherwise the change is "within the margin". Within a series, a different size or design (a mixture of experts that runs only part of the model, or a dense one) also affects the difference.
  • New series: a family not yet listed joins automatically when it passes a 1% share, its developer publishes the original weights, and it is not a fine-tune of another series.