Everything measured on one Mac Studio, the mothership
On a Mac Studio (M5 Max, 36GB),how well
can it translate?
- Memory
- 36GB
- Chip
- M5 Max
- CPU
- 18 cores
- GPU
- 32 cores
The mothership: Mac Studio (model MHL64J/A), macOS 27.0. Every number on this site was measured on this one Mac. Whether an AI runs at all comes down mostly to how much memory you have.
Local LLMs translate English sentences into Japanese and Japanese sentences into English, and the results are compared. You can translate on your own Mac without sending text over the internet. The focus is on English into Japanese.
Picks by use
Chosen from the scores and times below by fixed rules. Models within the margin are said to be level.
- If unsure, start here: gpt-oss:20b (The best English → Japanese within 5 seconds a sentence)
- Best English → Japanese: muse-glimmer:30b (The highest English → Japanese score)
- Best Japanese → English: muse-glimmer:30b (The highest Japanese → English score)
- Without the wait: gemma3:27b (The best English → Japanese within 2 seconds a sentence)
Translation scores
Highest English → Japanese score first. Bold marks the best value in each column; ± is the margin of error.
| Model | English → Japanese | Japanese → English | Time per sentence (EN→JA) | Not translated |
|---|---|---|---|---|
| muse-glimmer:30b | 36 pts | 60 pts | about 19 s | |
| gemma4:26b | 34 pts | 58 pts | about 8.5 s | |
| gpt-oss:20b | 33 pts | 55 pts | about 4.1 s | |
| gemma3:27b | 33 pts | 58 pts | about 1.6 s | |
| glm-4.7-flash:latest | 32 pts | 53 pts | about 23.3 s | |
| qwen3.8:27b | 31 pts | 58 pts | about 22 s | |
| qwen3-coder:30b | 31 pts | 55 pts | about 0.4 s | |
| qwen3.6:35b-a3b | 29 pts | 54 pts | about 28.9 s | |
| MiniCPM5-2B | 23 pts | 48 pts | about 9 s |
Compare on the same sentence
The same sentence as a professional translator put it (the reference) and as each model did. It shows what a score can't: misread meanings and differences in wording.
How it was measured
- Sentences: taken from a set of English news sentences with professional Japanese translations, leaving out headlines and very short or long lines. Each model translates 100 sentences each way, different ones for each direction.
- Score: how much the AI's translation overlaps with the reference in runs of the same characters, out of 100 (a measure called chrF). Different wording lowers it even when the meaning is right, so even a correct translation doesn't reach 100; use it to compare models. The ± shows how much the score moves when different sentences are picked.
- Time per sentence: the time to finish translating one sentence (the median). For models that think before answering, the thinking time is included.
- Not translated: sentences where the model thought too long and gave no translation within the set length, or started repeating itself and was stopped. They count as zero, as does a sentence whose reply Ollama couldn't read three times in a row.
- Everything is measured on the same Mac, each model with the settings its maker recommends (randomness turned off for a model with none; since October 2, 2026). The reference is one way to translate a sentence; other correct translations exist.
- To use it for translating: run the model in Ollama as for text AI and ask it to translate the text naturally into Japanese (how to try it).