For beginners
How to read
a model name
Names like "gemma-4-E4B-it", "30B-A3B" and "Q4_K_M" are a row of parts separated by hyphens. This guide explains each part (generation, size, tuning, quantization) with examples.
New to local LLMs? Start with getting started.
Example: gemma-4-E4B-it
gemmaName
(Google's series)4Generation
(the 4th)E4BSize
(effectively 4 billion)itTuning
(for chat)
Read the parts between the hyphens from the front. Which parts appear, and in what order, varies by model, but the kinds of parts are mostly the same.
Name and generation
| Written as | Meaning | Example |
|---|---|---|
gemma-4 / Qwen3.5 / LFM2.5 | The series and its generation. A bigger number is newer. | gemma-4-E4B-it / Qwen3.5-4B |
Nemotron-3-Nano-4B | With two numbers, the one with a B is the size; the other is the generation. | NVIDIA-Nemotron-3-Nano-4B-BF16 |
v2 | A revised release of the same model (version 2). | NVIDIA-Nemotron-Nano-12B-v2 |
Size
| Written as | Meaning | Example |
|---|---|---|
4B / 12b / 32B | The number of parameters (the parts of the AI's "brain"). B is billion: 4B is 4 billion; a small b is the same. Bigger tends to be smarter but uses more memory and runs slower. | Ornith-1.0-9B / DeepSeek-R1-Distill-Qwen-32B |
E4B | E is "effective". The file holds more parts, but it runs with the weight of about 4 billion. | gemma-4-E4B-it |
30B-A3B / 35b-a3b | 30 billion in all, but only 3 billion work at a time (A is active). It needs memory for all 30B but runs at nearly the speed of a 3B. This design is called MoE (mixture of experts). | NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 / qwen3.6:35b-a3b |
Nano / Mini / Small / Flash | Size or speed nicknames. What they mean in billions differs by maker, so go by the number. | Mistral-Small-3.2-24B-Instruct-2506 / glm-4.7-flash:latest |
Memory rule of thumb: about 0.7GB per billion at Q4_K_M (measured on this site), so about 6GB for 9B and 10GB for 14B. For MoE, count all the parameters. Check what fits your Mac on the picks by memory size.
Tuning and specialty
| Written as | Meaning | Example |
|---|---|---|
it / Instruct / Chat | Tuned to follow instructions and chat. This is the one you normally use. | gemma-4-E4B-it / Ministral-3-14B-Instruct-2512 |
Base | The raw version before tuning. It only continues text and is poor at chat. A name with no tag may be either, depending on the maker (Qwen adds Base to the raw one). | — |
Thinking / Reasoning / R1 | Thinks before answering. It writes out its reasoning first: stronger on hard problems, slower to reply. | DeepSeek-R1-0528-Qwen3-8B |
Distill | A small model taught to answer like a big one. "R1-Distill-Qwen-32B" is Qwen's 32B taught DeepSeek-R1's way of thinking. | DeepSeek-R1-Distill-Qwen-32B |
Coder | Tuned for programming. | qwen3-coder:30b |
JP | Tuned for Japanese. | LFM2.5-1.2B-JP-202606 |
VL / Vision | Reads images as well as text. | — |
Release date
| Written as | Meaning | Example |
|---|---|---|
2512 / 2506 | Year and month: 2512 is December 2025. | Ministral-3-14B-Instruct-2512 / Magistral-Small-2506 |
0414 / 0528 | Month and day: 0414 is April 14. Whether it is year-month or month-day depends on the maker. | GLM-4-9B-0414 / DeepSeek-R1-0528-Qwen3-8B |
202606 | Year and month in six digits: June 2026. | LFM2.5-1.2B-JP-202606 |
File format and quantization
| Written as | Meaning | Example |
|---|---|---|
GGUF | The file format Ollama and similar apps use. Almost every model on this site is in it. | gemma-4-12b-it |
MLX | A format for Apple silicon Macs (used by LM Studio and others). | — |
Q4_K_M | How much the model is compressed (quantized). Q4 squeezes each number to about 4 bits, K is the method, M is medium (S small, M medium, L large: how finely the important parts are kept). Smaller and faster, slightly less smart. This site's standard. | — |
Q8_0 | Squeezed to 8 bits. Close to the original, but about twice the size of Q4. | — |
BF16 / F16 | Not compressed (16 bits). About 3.5 times the size of Q4. | NVIDIA-Nemotron-3-Nano-4B-BF16 |
MXFP4 | Another 4-bit method. gpt-oss is released in it. | gpt-oss:20b |
Names in Ollama
| Written as | Meaning | Example |
|---|---|---|
gemma4:12b | Before the colon is the model, after it the tag (size or quantization). | gemma4:26b |
glm-4.7-flash:latest | latest is the recommended version, used when you leave the tag out. | glm-4.7-flash:latest |
A file from Hugging Face used in Ollama gets a longer name like this.
hf.coFrom
Hugging FaceunslothWho
shared itgemma-4-12b-it-GGUFRepository
(model + format)Q4_K_MWhich
quantization
Whoever shared it may not be the maker. Here unsloth converted Google's model to GGUF and shared it.
What a name does not tell you
- Even at about the same size (7-10 billion), Ornith-1.0-9B gets 83% of the knowledge questions right and GLM-4-9B-0414 gets 66%.
- Nor does a name tell you reply time, how good it is in Japanese, or how well it translates. See the measured results on the text page and the translation page.