open-source · 37 models · Apple + NVIDIA

Pick the right AI stack.
Run it on your machine.

Fingerprint your hardware, get the open-weight model + quant that actually fits, and install it in one command. No more CUDA/Metal/VRAM guesswork.

$npx runlocal-sh
$curl -fsSL https://runlocal.sh/install.sh | sh

Will it run?

Pick your machine, name a model — get the verdict instantly.

runs the same engine as the CLI

Popular: DeepSeek-R1 (distill) on Apple M2 Pro Llama 3.3 on NVIDIA GeForce RTX 4090 Qwen2.5-Coder on Apple M3 Max Gemma 3 on Apple M1

Best for your machine: Apple M1 · 8GB Apple M2 Pro · 32GB Apple M4 Pro · 48GB NVIDIA GeForce RTX 4090 · 24GB VRAM

Or ask the catalog

architecture kin · use cases · what fits your rig — no LLM

How it works

1 · Fingerprint
Detects your chip/GPU, unified memory or VRAM, bandwidth, and installed runtimes (Ollama, llama.cpp, MLX).
2 · Fit + speed math
GQA-aware KV-cache + quant sizing against your real memory budget → an honest "comfortable / tight / won't fit" verdict and a tok/s range.
3 · One-command install
runlocal install <model> pulls the best-fitting quant via Ollama and you're chatting.

Want the actual math — memory budgets, GQA KV-cache, MBU calibration? Read the full methodology →

Browse models

37 curated open-weight models, each with accurate fit data. See the full catalog →

Or take in the whole landscape at once: the compatibility matrix → — every model × every machine, one page.