A general chat model for everyday local conversation. The smallest cataloged machine that runs Gemma 3 is a Apple M1 · 8GB. Ranked below, smallest memory first — each model at its best quant, one model loaded at a time.
| machine | Gemma 3 |
|---|---|
| Apple M1 · 8GB | 1B @ Q8_0 · ~35–55 tok/s |
| Apple M2 · 8GB | 1B @ Q8_0 · ~50–80 tok/s |
| Apple M3 · 8GB | 1B @ Q8_0 · ~50–80 tok/s |
| AMD Radeon RX 7600 · 8GB VRAM | 4B @ Q8_0 · ~20–35 tok/s |
| NVIDIA GeForce RTX 3070 · 8GB VRAM | 4B @ Q8_0 · ~30–55 tok/s |
| NVIDIA GeForce RTX 4060 · 8GB VRAM | 4B @ Q8_0 · ~20–35 tok/s |
| NVIDIA GeForce RTX 3080 · 10GB VRAM | 4B @ Q8_0 · ~60–95 tok/s |
| AMD Radeon RX 6700 XT · 12GB VRAM | 12B @ Q4_K_M · ~14–25 tok/s |
| AMD Radeon RX 7700 XT · 12GB VRAM | 12B @ Q4_K_M · ~16–25 tok/s |
| NVIDIA GeForce RTX 3060 · 12GB VRAM | 12B @ Q4_K_M · ~14–25 tok/s |
| NVIDIA GeForce RTX 4070 · 12GB VRAM | 12B @ Q4_K_M · ~20–35 tok/s |
| Apple M1 · 16GB | 4B @ Q8_0 · ~7–12 tok/s |
Fits computed by the same engine as the CLI — reproduce this table with:
Already have a machine? Check your exact rig →