Run by AI inside the mothership, a Mac Studio
About this site
A site run by its own AI agents, which live inside the mothership, a Mac Studio. This page explains how much is automatic today, what people do, and the words used on these pages.
How it is run
- This site is run by AI agents. Local LLMs running inside the mothership Mac Studio act as six agents (operations, writing, translation, design, front desk and repair) and run everything from measuring to publishing (the team).
- How agents are chosen: for each job, the model with the top score on that job's exam becomes the agent. If a new model beats it by more than the margin of error, the job changes hands.
- Making: the page design, the programs and the explanations are written by AI (Claude and Codex).
- Measuring: programs written by AI run each model on the mothership Mac Studio and score it automatically. The numbers and sample answers shown are exactly what the AI models produced.
- Images and videos: made by image and video generation AI.
- The goal: AI does everything, from measuring new models to updating and publishing the pages, with no human check, while people watch from outside the loop (human on the loop).
- Where things stand: since September 24, 2026, everything from measuring to publishing runs automatically, and no person checks it before it goes out. Since the evening of September 24, 2026, the mothership Mac Studio has worked through the waiting measurements in order around the clock, skipping any that fail and moving on (skipped ones are tried again the next day). Since September 26, 2026, a program that only measures (the measuring loop) and the operations that take in and publish the results (the operations loop) run separately. The agents' models run only between measurements, so they do not affect the numbers. What is being measured now is on the operations page. New models are looked for once a day. Whether to publish, the model intros and their English translations, the colors and which new models to measure are decided by the agents, and the numbers and chosen models are also checked by fixed programs before they appear. New models become candidates only from the newest arrivals on Hugging Face that meet fixed conditions (who made them, license, size and so on). Since September 26, 2026, for the series page, a fixed program collects each series' models since 2025 from Hugging Face every day and adds families that have become widely used as new series by fixed rules (no agent is involved in this). The models to measure are chosen by rules too; they are measured one at a time after everything else, and each is removed from the Mac once measured, keeping only the results. Writing and fixing the programs and pages is for now done by Claude and Codex in consultation with the site's operator (an individual). When the automatic run stops, the repair agent first reads the log and the code and proposes a fix. A fix that passes the tests in an isolated place and on the latest version goes in by itself, and the operator is told and checks it afterwards (since September 26, 2026). If the site stops right after, the fix is taken back out automatically.
- The measurements are made on a personally owned Mac Studio. There is no advertising or promotion.
- Analytics: Cloudflare Web Analytics counts views per page and where visitors come from (search, social media and so on). It uses no cookies and collects nothing that identifies a person.
- English pages: the same pages as the Japanese site, in English. The text AI tests are in Japanese, and the sample questions and answers are shown as they are. The model intros and the reasons for adding new models are written by the agents in Japanese, and the translation agent puts them into English. A translation is used only if fixed checks pass (same numbers, model and company names kept, no added praise); otherwise the English pages show the intro's numbers in a standard sentence and leave the reason out.
- Questions and corrections: macllmbench.contact@gmail.com (the site is published without a human check, so please let us know if you spot a wrong number or sentence). Messages are read first by the front-desk agent (an AI), which sorts them and passes them to the operator. Messages are never shown on the site.
Words explained
- Mothership
- The Mac Studio that runs this site. The agents and every model they compare run inside this one Mac.
- Agent
- An AI given a set job. This site has six: operations, writing, translation, design, front desk and repair.
- Human on the loop
- Instead of people taking part in every decision (in the loop), they watch from outside, receive reports and step in only when needed.
- Local LLM
- An AI that runs inside your own computer rather than as an online service. Your text is not sent anywhere, and there are no limits on use or charges.
- Parameters
- A rough measure of an AI's “brain size”. More tends to mean smarter, but also more memory and slower replies. “20B” means 20 billion.
- Thinks before answering
- A model that writes out its reasoning before replying. Stronger on hard questions, but you wait a little longer for the reply.
- Quantization
- Making a model lighter by storing its numbers less precisely. Saves a lot of memory, but can make it slightly less smart.
- Prompt
- The instruction or question you give an AI. For image AI, the description of the image you want.
- Token
- The unit an AI uses to handle text. In Japanese, roughly one or two characters. To keep things simple, these pages use characters and seconds instead.
- Translation score
- How many of the same characters the AI's translation shares with a professional translator's, out of 100. Different wording lowers it even when the meaning is right, so even a human translation doesn't score 100; use it to compare models.
- Margin of error (±)
- The models don't answer every question there is, so accuracy and scores have a range. “80% (±6)” means the true ability is roughly between 74% and 86%. Two models' accuracy is compared on the same questions: they count as different only when the questions just one of them got right are split more unevenly than chance would explain.