Best local AI for NVIDIA GeForce RTX 5090 · 32GB VRAM

The best local AI on a NVIDIA GeForce RTX 5090 · 32GB VRAM right now is DeepSeek-R1 (distill) (32B @ Q3_K_M, ~45–70 tok/s). Full ranking below — capability × real speed on this hardware, never sponsored.

#modelbest fitverdictspeed
1 DeepSeek-R1 (distill) 32B Q3_K_M ✅ Runs comfortably 45–70 tok/s
2 Qwen2.5 32B Q3_K_M ✅ Runs comfortably 45–70 tok/s
3 Qwen3 32B Q3_K_M ✅ Runs comfortably 45–70 tok/s
4 Llama 3.1 70B Q2_K ✅ Runs (tight) 25–40 tok/s
5 Gemma 3 27B Q3_K_M ✅ Runs comfortably 45–75 tok/s
6 Llama 3.3 70B Q2_K ✅ Runs (tight) 25–40 tok/s
7 Mistral Small 3 24B Q5_K_M ✅ Runs comfortably 45–70 tok/s
8 Gemma 2 27B Q4_K_M ✅ Runs comfortably 40–65 tok/s
9 Gemma 4 12B Q8_0 ✅ Runs comfortably 50–80 tok/s
10 Mixtral 8x7B 8x7B (MoE) Q4_K_M ✅ Runs (tight) 85–140 tok/s
11 Yi 1.5 34B Q3_K_M ✅ Runs comfortably 40–70 tok/s
12 Phi-4 14B Q8_0 ✅ Runs comfortably 45–75 tok/s
$npx runlocal-sh install deepseek-r1

Rankings use the same GQA-aware fit + bandwidth-calibrated speed math as the CLI — no sponsorships, no affiliate re-sorting, ever.

Different machine? Apple M1 · 8GB · Apple M1 Pro · 16GB · Apple M2 · 16GB · Apple M3 Pro · 18GB · Apple M2 Pro · 32GB · Apple M4 Pro · 48GB

Or check a specific model on your exact rig: open the checker →