Hardware-to-Model Matchmaker
Pick your GPU — or enter your VRAM — and see exactly which open-weight models you can self-host, at which quantization, with how much headroom. All models listed are fully sovereign: open weights, self-hosted, zero external API calls.
Hardware and workload
Models that fit your hardware
Best (least lossy) quantization that fits is shown. Larger models are listed first.
| Model | License | Context | Quantization | Needs | Headroom | Runtimes |
|---|---|---|---|---|---|---|
Gemma 2 27B | Gemma Terms of Use | 8K | Q4 | 22 GB | 8% | Ollama · vLLM |
StarCoder2 15B | BigCode OpenRAIL-M | 16K | INT8 | 22 GB | 9% | Ollama · vLLM |
Code Llama 13B | Llama 2 Community License | 16K | INT8 | 20 GB | 19% | Ollama · vLLM |
Mistral Nemo 12B Long-context document intelligence at mid size. | Apache-2.0 | 128K | INT8 | 19 GB | 24% | Ollama · vLLM |
Gemma 2 9B Strong reasoning; short 8K context. | Gemma Terms of Use | 8K | INT8 | 15 GB | 39% | Ollama · vLLM |
Llama 3.1 8B The default starting point for private RAG on a single GPU. | Llama 3.1 Community License | 128K | FP16 | 23 GB | 6% | Ollama · vLLM · OpenShift AI |
Mistral 7B Efficient permissively-licensed baseline. | Apache-2.0 | 32K | FP16 | 21 GB | 15% | Ollama · vLLM · Docker |
Qwen 2.5 7B | Apache-2.0 | 128K | FP16 | 21 GB | 15% | Ollama · vLLM |
Phi-3.5 Mini (3.8B) Runs on laptops and edge devices; excellent quality-to-size ratio. | MIT | 128K | FP16 | 13 GB | 46% | Ollama · Edge · Docker |
Needs more memory
Minimum requirement shown at the smallest common quantization with your context and concurrency.
DeepSeek R1 Distill 32B
Needs ≈ 26 GB at Q4
Mixtral 8×7B
Needs ≈ 36 GB at Q4
Llama 3.1 70B
Needs ≈ 51 GB at Q4
Qwen 2.5 72B
Needs ≈ 52 GB at Q4
Llama 3.1 405B
Needs ≈ 274 GB at Q4
DeepSeek R1 (671B MoE)
Needs ≈ 451 GB at Q4
Don't own the hardware yet? Rent it by the hour first.
Before buying GPUs, validate your model choice on rented hardware. Run your actual workload for a few dollars, measure tokens/sec, then size your on-prem purchase with real numbers.
Email me these results + sovereign AI deployment notes
Get your hardware match summary and occasional practical notes on self-hosting LLMs, private RAG, and air-gapped deployment. No spam.
Hardware Sizing Guide
Work the estimate the other way: start from a model class and size the hardware you need.
Model Registry
Sovereignty tiers, deployment runtimes, and a decision guide for private AI model selection.
Planning a real deployment?
Get a fixed-scope architecture review: model selection, hardware sizing, runtime topology, and a written findings document for your private RAG or air-gapped AI platform.
Estimates are approximate. Actual memory depends on runtime, batching, quantization format, KV cache implementation, and model architecture. Always benchmark before purchasing hardware.