A chat model plus an embedder for retrieval over your own documents. The smallest cataloged machine that runs Llama 3.3 + Nomic Embed Text is a NVIDIA GeForce RTX 5090 · 32GB VRAM. Ranked below, smallest memory first — each model at its best quant, one model loaded at a time.
| machine | Llama 3.3 | Nomic Embed Text |
|---|---|---|
| NVIDIA GeForce RTX 5090 · 32GB VRAM | 70B @ Q2_K · ~25–40 tok/s | 137M @ F16 · ~1350–2260 tok/s |
| AMD Radeon Pro W7900 · 48GB VRAM | 70B @ Q3_K_M · ~9–15 tok/s | 137M @ F16 · ~620–1030 tok/s |
| Apple M3 Max · 48GB | 70B @ Q2_K · ~6–10 tok/s | 137M @ F16 · ~330–550 tok/s |
| Apple M4 Max · 48GB | 70B @ Q2_K · ~8–13 tok/s | 137M @ F16 · ~410–690 tok/s |
| Apple M4 Pro · 48GB | 70B @ Q2_K · ~5–8 tok/s | 137M @ F16 · ~250–420 tok/s |
| Apple M1 Max · 64GB | 70B @ Q4_K_M · ~4–7 tok/s | 137M @ F16 · ~330–550 tok/s |
| Apple M1 Ultra · 64GB | 70B @ Q4_K_M · ~5–9 tok/s | 137M @ F16 · ~440–730 tok/s |
| Apple M2 Max · 64GB | 70B @ Q4_K_M · ~4–7 tok/s | 137M @ F16 · ~330–550 tok/s |
| Apple M2 Ultra · 64GB | 70B @ Q4_K_M · ~5–9 tok/s | 137M @ F16 · ~440–730 tok/s |
| Apple M2 Max · 96GB | 70B @ Q5_K_M · ~4–6 tok/s | 137M @ F16 · ~330–550 tok/s |
| Apple M2 Ultra · 128GB | 70B @ Q8_0 · ~3–5 tok/s | 137M @ F16 · ~440–730 tok/s |
| Apple M3 Max · 128GB | 70B @ Q8_0 · ~2–4 tok/s | 137M @ F16 · ~330–550 tok/s |
Fits computed by the same engine as the CLI — reproduce this table with:
Already have a machine? Check your exact rig →