The best local AI on a Apple M1 · 8GB right now is DeepSeek-R1 (distill) (1.5B @ Q4_K_M, ~30–50 tok/s). Full ranking below — capability × real speed on this hardware, never sponsored.
| # | model | best fit | verdict | speed |
|---|---|---|---|---|
| 1 | DeepSeek-R1 (distill) | 1.5B Q4_K_M | ✅ Runs comfortably | 30–50 tok/s |
| 2 | Qwen2.5 | 3B Q5_K_M | ✅ Runs (tight) | 16–25 tok/s |
| 3 | Qwen3 | 1.7B Q4_K_M | ✅ Runs (tight) | 20–35 tok/s |
| 4 | Llama 3.2 | 1B Q4_K_M | ✅ Runs comfortably | 40–65 tok/s |
| 5 | Gemma 2 | 2B Q4_K_M | ✅ Runs (tight) | 16–25 tok/s |
| 6 | Gemma 3 | 1B Q4_K_M | ✅ Runs comfortably | 50–85 tok/s |
| 7 | TinyLlama | 1.1B Q4_K_M | ✅ Runs comfortably | 50–85 tok/s |
| 8 | Qwen2.5-Coder | 1.5B Q8_0 | ✅ Runs comfortably | 25–40 tok/s |
| 9 | SmolLM3 3B | 3B Q4_K_M | ✅ Runs (tight) | 16–25 tok/s |
| 10 | Nomic Embed Text | 137M F16 | ✅ Runs comfortably | 75–120 tok/s |
| 11 | StarCoder2 | 3B Q4_K_M | ✅ Runs (tight) | 19–30 tok/s |
| 12 | mxbai-embed-large | 335M F16 | ✅ Runs comfortably | 30–50 tok/s |
Rankings use the same GQA-aware fit + bandwidth-calibrated speed math as the CLI — no sponsorships, no affiliate re-sorting, ever.
Different machine? Apple M1 Pro · 16GB · Apple M2 · 16GB · Apple M3 Pro · 18GB · Apple M2 Pro · 32GB · Apple M4 Pro · 48GB · Apple M3 Max · 36GB
Or check a specific model on your exact rig: open the checker →