Everything measured on one Mac Studio, the mothership

On a Mac Studio (M5 Max, 36GB),how well
can it translate?

Memory
36GB
Chip
M5 Max
CPU
18 cores
GPU
32 cores

The mothership: Mac Studio (model MHL64J/A), macOS 27.0. Every number on this site was measured on this one Mac. Whether an AI runs at all comes down mostly to how much memory you have.

Local LLMs translate English sentences into Japanese and Japanese sentences into English, and the results are compared. You can translate on your own Mac without sending text over the internet. The focus is on English into Japanese.

Picks by use

Chosen from the scores and times below by fixed rules. Models within the margin are said to be level.

  • If unsure, start here: gpt-oss:20b (The best English → Japanese within 5 seconds a sentence)
  • Best English → Japanese: muse-glimmer:30b (The highest English → Japanese score)
  • Best Japanese → English: muse-glimmer:30b (The highest Japanese → English score)
  • Without the wait: gemma3:27b (The best English → Japanese within 2 seconds a sentence)

Translation scores

Highest English → Japanese score first. Bold marks the best value in each column; ± is the margin of error.

Model English → Japanese Japanese → English Time per sentence (EN→JA) Not translated
muse-glimmer:30bMeta36 pts60 ptsabout 19 s
gemma4:26bGoogle34 pts58 ptsabout 8.5 s
gpt-oss:20bOpenAI33 pts55 ptsabout 4.1 s
gemma3:27bGoogle33 pts58 ptsabout 1.6 s
glm-4.7-flash:latestZ.ai32 pts53 ptsabout 23.3 s
qwen3.8:27bAlibaba31 pts58 ptsabout 22 s
qwen3-coder:30bAlibaba31 pts55 ptsabout 0.4 s
qwen3.6:35b-a3bAlibaba29 pts54 ptsabout 28.9 s
MiniCPM5-2BOpenBMB23 pts48 ptsabout 9 s

Compare on the same sentence

The same sentence as a professional translator put it (the reference) and as each model did. It shows what a score can't: misread meanings and differences in wording.

How it was measured

  • Sentences: taken from a set of English news sentences with professional Japanese translations, leaving out headlines and very short or long lines. Each model translates 100 sentences each way, different ones for each direction.
  • Score: how much the AI's translation overlaps with the reference in runs of the same characters, out of 100 (a measure called chrF). Different wording lowers it even when the meaning is right, so even a correct translation doesn't reach 100; use it to compare models. The ± shows how much the score moves when different sentences are picked.
  • Time per sentence: the time to finish translating one sentence (the median). For models that think before answering, the thinking time is included.
  • Not translated: sentences where the model thought too long and gave no translation within the set length, or started repeating itself and was stopped. They count as zero, as does a sentence whose reply Ollama couldn't read three times in a row.
  • Everything is measured on the same Mac, each model with the settings its maker recommends (randomness turned off for a model with none; since October 2, 2026). The reference is one way to translate a sentence; other correct translations exist.
  • To use it for translating: run the model in Ollama as for text AI and ask it to translate the text naturally into Japanese (how to try it).