What hardware runs a local reasoning workflow?

A chain-of-thought reasoning model for harder problems. The smallest cataloged machine that runs DeepSeek-R1 (distill) is a AMD Radeon RX 7600 · 8GB VRAM. Ranked below, smallest memory first — each model at its best quant, one model loaded at a time.

machineDeepSeek-R1 (distill)
AMD Radeon RX 7600 · 8GB VRAM 8B (Llama) @ Q5_K_M · ~17–30 tok/s
Apple M1 · 8GB 1.5B @ Q8_0 · ~19–30 tok/s
Apple M2 · 8GB 1.5B @ Q8_0 · ~30–45 tok/s
Apple M3 · 8GB 1.5B @ Q8_0 · ~30–45 tok/s
NVIDIA GeForce RTX 3070 · 8GB VRAM 8B (Llama) @ Q5_K_M · ~25–45 tok/s
NVIDIA GeForce RTX 4060 · 8GB VRAM 8B (Llama) @ Q5_K_M · ~17–30 tok/s
NVIDIA GeForce RTX 3080 · 10GB VRAM 14B @ Q3_K_M · ~35–60 tok/s
AMD Radeon RX 6700 XT · 12GB VRAM 14B @ Q4_K_M · ~14–25 tok/s
AMD Radeon RX 7700 XT · 12GB VRAM 14B @ Q4_K_M · ~16–25 tok/s
NVIDIA GeForce RTX 3060 · 12GB VRAM 14B @ Q4_K_M · ~14–25 tok/s
NVIDIA GeForce RTX 4070 · 12GB VRAM 14B @ Q4_K_M · ~20–35 tok/s
AMD Radeon RX 6800 XT · 16GB VRAM 14B @ Q5_K_M · ~17–30 tok/s

Fits computed by the same engine as the CLI — reproduce this table with:

$npx runlocal-sh advise deepseek-r1

Already have a machine? Check your exact rig →