Fast, high-quality text embeddings for RAG. Tiny — runs anywhere.
embeddingsrag · sizes: 137M · Apache-2.0
| machine | best fit | verdict | speed |
|---|---|---|---|
| Apple M1 · 8GB | 137M F16 | ✅ Runs comfortably | 75–120 tok/s |
| Apple M1 Pro · 16GB | 137M F16 | ✅ Runs comfortably | 200–330 tok/s |
| Apple M2 · 16GB | 137M F16 | ✅ Runs comfortably | 110–180 tok/s |
| Apple M3 Pro · 18GB | 137M F16 | ✅ Runs comfortably | 150–250 tok/s |
| Apple M2 Pro · 32GB | 137M F16 | ✅ Runs comfortably | 200–330 tok/s |
| Apple M4 Pro · 48GB | 137M F16 | ✅ Runs comfortably | 250–420 tok/s |
| Apple M3 Max · 36GB | 137M F16 | ✅ Runs comfortably | 330–550 tok/s |
| Apple M2 Max · 64GB | 137M F16 | ✅ Runs comfortably | 330–550 tok/s |
| Apple M4 Max · 128GB | 137M F16 | ✅ Runs comfortably | 410–690 tok/s |
| Apple M2 Ultra · 192GB | 137M F16 | ✅ Runs comfortably | 440–730 tok/s |
| NVIDIA GeForce RTX 3060 · 12GB VRAM | 137M F16 | ✅ Runs comfortably | 270–450 tok/s |
| NVIDIA GeForce RTX 4070 · 12GB VRAM | 137M F16 | ✅ Runs comfortably | 380–630 tok/s |
| NVIDIA GeForce RTX 4090 · 24GB VRAM | 137M F16 | ✅ Runs comfortably | 760–1270 tok/s |
| NVIDIA GeForce RTX 5080 · 16GB VRAM | 137M F16 | ✅ Runs comfortably | 730–1210 tok/s |
| NVIDIA GeForce RTX 5090 · 32GB VRAM | 137M F16 | ✅ Runs comfortably | 1350–2260 tok/s |
| AMD Radeon RX 7900 XTX · 24GB VRAM | 137M F16 | ✅ Runs comfortably | 690–1150 tok/s |
Shopping for a machine? What hardware do I need to run Nomic Embed Text? →
Check your exact rig: open the checker →