What hardware runs a local local RAG workflow?

A chat model plus an embedder for retrieval over your own documents. The smallest cataloged machine that runs Llama 3.3 + Nomic Embed Text is a NVIDIA GeForce RTX 5090 · 32GB VRAM. Ranked below, smallest memory first — each model at its best quant, one model loaded at a time.

machineLlama 3.3Nomic Embed Text
NVIDIA GeForce RTX 5090 · 32GB VRAM 70B @ Q2_K · ~25–40 tok/s 137M @ F16 · ~1350–2260 tok/s
AMD Radeon Pro W7900 · 48GB VRAM 70B @ Q3_K_M · ~9–15 tok/s 137M @ F16 · ~620–1030 tok/s
Apple M3 Max · 48GB 70B @ Q2_K · ~6–10 tok/s 137M @ F16 · ~330–550 tok/s
Apple M4 Max · 48GB 70B @ Q2_K · ~8–13 tok/s 137M @ F16 · ~410–690 tok/s
Apple M4 Pro · 48GB 70B @ Q2_K · ~5–8 tok/s 137M @ F16 · ~250–420 tok/s
Apple M1 Max · 64GB 70B @ Q4_K_M · ~4–7 tok/s 137M @ F16 · ~330–550 tok/s
Apple M1 Ultra · 64GB 70B @ Q4_K_M · ~5–9 tok/s 137M @ F16 · ~440–730 tok/s
Apple M2 Max · 64GB 70B @ Q4_K_M · ~4–7 tok/s 137M @ F16 · ~330–550 tok/s
Apple M2 Ultra · 64GB 70B @ Q4_K_M · ~5–9 tok/s 137M @ F16 · ~440–730 tok/s
Apple M2 Max · 96GB 70B @ Q5_K_M · ~4–6 tok/s 137M @ F16 · ~330–550 tok/s
Apple M2 Ultra · 128GB 70B @ Q8_0 · ~3–5 tok/s 137M @ F16 · ~440–730 tok/s
Apple M3 Max · 128GB 70B @ Q8_0 · ~2–4 tok/s 137M @ F16 · ~330–550 tok/s

Fits computed by the same engine as the CLI — reproduce this table with:

$npx runlocal-sh advise llama-3.3, nomic-embed-text

Already have a machine? Check your exact rig →