Apple LLM Performance Tracker
Which open-weight models actually run well on which Apple silicon, and through which engine — per CPU, memory and machine count, from M1 to M5 Ultra.
The problem, in one sentence
The model cards say what a model needs; they do not say what your Mac can do with it. A 27B dense model and a 284B MoE model can both nominally fit in 256 GB, but only one of them is actually fast there — because the difference is bandwidth, not capacity, and the right answer changes with the engine you load it through.
What the page does
- Cluster picker. Every Apple silicon generation from M1 through M5 Ultra, 96 GB to 512 GB per machine, up to six machines pooled. The pooled arithmetic is explicit: 2 × M5 Ultra is 512 GB pooled, and a model spread across machines still pays the Thunderbolt hop on every token.
- One row per model. Status (runs / runs degraded / purpose-built / too large), the best engine for that cluster, the build the engine should load at that memory size, and the free memory after the weights are resident.
- One card per model. Architecture, licence, context, a written verdict, the quant ladder with measured bits-per-weight, the fidelity band (with the published evidence behind each threshold), the per-token KV cost derived from the model’s own
config.json, and the tracked upstream issues — Apple-silicon ones only. - Ten use-case categories (agentic, coding, terminal work, image generation, TTS, …) with curated rankings. The rankings are explicit lists with evidence, not a computed score — Terminal-Bench 2.0 and Terminal-Bench 2.1 are different tests, and the page refuses to rank across them.
Engines covered
| Engine | Format | Interface |
|---|---|---|
| llama.cpp | GGUF | CLI + llama-server |
| Ollama | MLX on Apple silicon, GGUF elsewhere | background server + CLI |
| LM Studio | GGUF and MLX | desktop app + lms + server |
| oMLX | MLX | menu-bar app + server |
| vLLM Metal | MLX | CLI + the vLLM server |
| vllm-mlx | MLX | server |
| mlx-lm | MLX | CLI + mlx_lm.server |
| DwarfStar / ds4 | purpose-built GGUF | CLI + server + agent |
Plus a separate set for generative media — mflux, MLX-Audio, MLX-Video, DiffusionKit — because none of the text servers can load a diffusion or audio model.
How the numbers are derived
- Weights are summed file sizes from the linked Hugging Face repositories — measured, not estimated.
- Bits per weight is
bytes * 8 / total parameters, the effective figure rather than whatever the quant is named. Those diverge badly on MoE models: GLM-5.2’sUD-IQ1_Sis really 2.33 bpw because the non-expert tensors are carried at higher precision. - KV cost per token counts only the layers whose cache actually grows with context — sliding-window layers are bounded by the window, and linear/Mamba layers hold a fixed recurrent state, so neither belongs in a per-token figure.
- Fit assumes a 90% wired-memory limit plus framework overhead (~10 GB for an LLM server, ~1.5 GB for an image or audio runtime). It answers “does this load”, not “does this run well”.
Nothing on the page has been benchmarked on the hardware by us: benchmark scores are vendor- or aggregator-reported, issue states are a twice-daily snapshot, and every memory figure is arithmetic over published specifications. Every number links to where it came from — that is the whole point of the project.
The code
Standard-library Python only, no dependencies. The data is one record, one file (data/models/<id>.py, data/engines/<id>.py, …), so two people adding two different models never touch the same file. tracker/validate.py machine-checks every record, tracker/build.py renders the single-file page, and a twice-daily workflow re-polls the tracked issue states from GitHub and deploys the change.
Fork and lineage
This project is a fork of dreamingwell/apple-llm-performance, the upstream that first assembled the model, engine and use-case records. It is kept in sync with upstream and is being slowly adapted to its own purpose as it diverges. The canonical home is the Pages site served from this repository; the page’s Open Graph and canonical tags, and the “Open source” link in its header, point here rather than at the upstream author’s Pages site.
- Live page (canonical): coltonkawamura.github.io/apple-llm-performance
- Fork repository: github.com/ColtonKawamura/apple-llm-performance
- Upstream: github.com/dreamingwell/apple-llm-performance