For beginners

How to read
a model name

Names like "gemma-4-E4B-it", "30B-A3B" and "Q4_K_M" are a row of parts separated by hyphens. This guide explains each part (generation, size, tuning, quantization) with examples.

New to local LLMs? Start with getting started.

Example: gemma-4-E4B-it

  1. gemmaName
    (Google's series)
  2. 4Generation
    (the 4th)
  3. E4BSize
    (effectively 4 billion)
  4. itTuning
    (for chat)

Read the parts between the hyphens from the front. Which parts appear, and in what order, varies by model, but the kinds of parts are mostly the same.

Name and generation

Written asMeaningExample
gemma-4 / Qwen3.5 / LFM2.5The series and its generation. A bigger number is newer.gemma-4-E4B-it / Qwen3.5-4B
Nemotron-3-Nano-4BWith two numbers, the one with a B is the size; the other is the generation.NVIDIA-Nemotron-3-Nano-4B-BF16
v2A revised release of the same model (version 2).NVIDIA-Nemotron-Nano-12B-v2

Size

Written asMeaningExample
4B / 12b / 32BThe number of parameters (the parts of the AI's "brain"). B is billion: 4B is 4 billion; a small b is the same. Bigger tends to be smarter but uses more memory and runs slower.Ornith-1.0-9B / DeepSeek-R1-Distill-Qwen-32B
E4BE is "effective". The file holds more parts, but it runs with the weight of about 4 billion.gemma-4-E4B-it
30B-A3B / 35b-a3b30 billion in all, but only 3 billion work at a time (A is active). It needs memory for all 30B but runs at nearly the speed of a 3B. This design is called MoE (mixture of experts).NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 / qwen3.6:35b-a3b
Nano / Mini / Small / FlashSize or speed nicknames. What they mean in billions differs by maker, so go by the number.Mistral-Small-3.2-24B-Instruct-2506 / glm-4.7-flash:latest

Memory rule of thumb: about 0.7GB per billion at Q4_K_M (measured on this site), so about 6GB for 9B and 10GB for 14B. For MoE, count all the parameters. Check what fits your Mac on the picks by memory size.

Tuning and specialty

Written asMeaningExample
it / Instruct / ChatTuned to follow instructions and chat. This is the one you normally use.gemma-4-E4B-it / Ministral-3-14B-Instruct-2512
BaseThe raw version before tuning. It only continues text and is poor at chat. A name with no tag may be either, depending on the maker (Qwen adds Base to the raw one).—
Thinking / Reasoning / R1Thinks before answering. It writes out its reasoning first: stronger on hard problems, slower to reply.DeepSeek-R1-0528-Qwen3-8B
DistillA small model taught to answer like a big one. "R1-Distill-Qwen-32B" is Qwen's 32B taught DeepSeek-R1's way of thinking.DeepSeek-R1-Distill-Qwen-32B
CoderTuned for programming.qwen3-coder:30b
JPTuned for Japanese.LFM2.5-1.2B-JP-202606
VL / VisionReads images as well as text.—

Release date

Written asMeaningExample
2512 / 2506Year and month: 2512 is December 2025.Ministral-3-14B-Instruct-2512 / Magistral-Small-2506
0414 / 0528Month and day: 0414 is April 14. Whether it is year-month or month-day depends on the maker.GLM-4-9B-0414 / DeepSeek-R1-0528-Qwen3-8B
202606Year and month in six digits: June 2026.LFM2.5-1.2B-JP-202606

File format and quantization

Written asMeaningExample
GGUFThe file format Ollama and similar apps use. Almost every model on this site is in it.gemma-4-12b-it
MLXA format for Apple silicon Macs (used by LM Studio and others).—
Q4_K_MHow much the model is compressed (quantized). Q4 squeezes each number to about 4 bits, K is the method, M is medium (S small, M medium, L large: how finely the important parts are kept). Smaller and faster, slightly less smart. This site's standard.—
Q8_0Squeezed to 8 bits. Close to the original, but about twice the size of Q4.—
BF16 / F16Not compressed (16 bits). About 3.5 times the size of Q4.NVIDIA-Nemotron-3-Nano-4B-BF16
MXFP4Another 4-bit method. gpt-oss is released in it.gpt-oss:20b

Names in Ollama

Written asMeaningExample
gemma4:12bBefore the colon is the model, after it the tag (size or quantization).gemma4:26b
glm-4.7-flash:latestlatest is the recommended version, used when you leave the tag out.glm-4.7-flash:latest

A file from Hugging Face used in Ollama gets a longer name like this.

  1. hf.coFrom
    Hugging Face
  2. unslothWho
    shared it
  3. gemma-4-12b-it-GGUFRepository
    (model + format)
  4. Q4_K_MWhich
    quantization

Whoever shared it may not be the maker. Here unsloth converted Google's model to GGUF and shared it.

What a name does not tell you