Free tool
Local-LLM GPU Checker
Pick your GPU (or type your VRAM) and see exactly which open LLMs you can run locally — at which quantization, with a fit verdict and a rough speed feel. Built on vendor-verified VRAM and real GGUF sizes.
RTX 4090 (24 GB) · Q4_K_M · 8K (typical chat)
You can comfortably run up to Gemma 3 27B at Q4_K_M.
- Qwen3 32B32.8B · QwenTight fitQ4_K_M ≈ 21.3 GB· ~39 tok/s (usable)
- Qwen2.5 32B32B · QwenTight fitQ4_K_M ≈ 20.8 GB· ~40 tok/s (fast)
- DeepSeek-R1 Distill 32B32B · DeepSeekTight fitQ4_K_M ≈ 20.8 GB· ~40 tok/s (fast)
- Gemma 3 27B27B · GemmaRuns wellQ4_K_M ≈ 17.5 GB· ~48 tok/s (fast)
- Mistral Small 3 24B24B · MistralRuns wellQ4_K_M ≈ 15.6 GB· ~53 tok/s (fast)
- Qwen3 14B14.8B · QwenRuns wellQ4_K_M ≈ 9.6 GB· ~87 tok/s (fast)
- Qwen2.5 14B14B · QwenRuns wellQ4_K_M ≈ 9.1 GB· ~92 tok/s (fast)
- Phi-4 14B14B · PhiRuns wellQ4_K_M ≈ 9.1 GB· ~92 tok/s (fast)
- DeepSeek-R1 Distill 14B14B · DeepSeekRuns wellQ4_K_M ≈ 9.1 GB· ~92 tok/s (fast)
- Gemma 3 12B12B · GemmaRuns wellQ4_K_M ≈ 7.8 GB· ~107 tok/s (fast)
- Gemma 2 9B9B · GemmaRuns wellQ4_K_M ≈ 5.8 GB· ~143 tok/s (fast)
- Qwen3 8B8.2B · QwenRuns wellQ4_K_M ≈ 5.3 GB· ~156 tok/s (fast)
- Llama 3.1 8B8B · LlamaRuns wellQ4_K_M ≈ 5.2 GB· ~160 tok/s (fast)
- DeepSeek-R1 Distill 8B8B · DeepSeekRuns wellQ4_K_M ≈ 5.2 GB· ~160 tok/s (fast)
- Mistral 7B7.3B · MistralRuns wellQ4_K_M ≈ 4.7 GB· ~160 tok/s (fast)
- Qwen2.5 7B7B · QwenRuns wellQ4_K_M ≈ 4.5 GB· ~160 tok/s (fast)
- Qwen3 4B4B · QwenRuns wellQ4_K_M ≈ 2.6 GB· ~160 tok/s (fast)
- Gemma 3 4B4B · GemmaRuns wellQ4_K_M ≈ 2.6 GB· ~160 tok/s (fast)
- Phi-3.5 mini (3.8B)3.8B · PhiRuns wellQ4_K_M ≈ 2.5 GB· ~160 tok/s (fast)
- Llama 3.2 3B3.2B · LlamaRuns wellQ4_K_M ≈ 2.1 GB· ~160 tok/s (fast)
- Gemma 2 2B2.6B · GemmaRuns wellQ4_K_M ≈ 1.7 GB· ~160 tok/s (fast)
- Qwen3 1.7B1.7B · QwenRuns wellQ4_K_M ≈ 1.1 GB· ~160 tok/s (fast)
- Qwen2.5 1.5B1.5B · QwenRuns wellQ4_K_M ≈ 1 GB· ~160 tok/s (fast)
- Llama 3.2 1B1.2B · LlamaRuns wellQ4_K_M ≈ 0.8 GB· ~160 tok/s (fast)
- Gemma 3 1B1B · GemmaRuns wellQ4_K_M ≈ 0.6 GB· ~160 tok/s (fast)
- Qwen3 0.6B0.6B · QwenRuns wellQ4_K_M ≈ 0.4 GB· ~160 tok/s (fast)
- Qwen2.5 72B72B · QwenWon't fitQ4_K_M ≈ 46.7 GB
- Llama 3.3 70B70B · LlamaWon't fitQ4_K_M ≈ 45.4 GB
- DeepSeek-R1 Distill 70B70B · DeepSeekWon't fitQ4_K_M ≈ 45.4 GB
Need a card for bigger models? See our GPU + mini-PC guides.
Estimates for DENSE models, anchored to real GGUF sizes: ≈ params × bytes-per-weight (Q3 ~0.43 → Q4 ~0.55 → Q8 ~1.06 → FP16 2.0) + a KV-cache factor for your context. Tokens/sec is a rough bandwidth-class “feel” (single user, short context) — not a benchmark; real speed depends on your engine (llama.cpp / Ollama / vLLM / MLX), batch size and context. Apple usable memory ≈ 72% of total. AMD (ROCm/Vulkan) + Intel (IPEX) ecosystems are less mature than CUDA. Re-verify before buying.
Frequently asked
How much VRAM do I need to run a model locally?
Rule of thumb: multiply the model's parameter count (in billions) by the bytes per weight for your quantization — 2 GB/B for FP16, 1 GB/B for 8-bit, ~0.5 GB/B for 4-bit (GGUF Q4_K_M) — then add about 20% for the KV cache, activations and framework overhead. So an 8B model at 4-bit needs roughly 8 × 0.5 × 1.2 ≈ 5 GB, and a 70B model at 4-bit needs about 42–46 GB.
Can I run a 70B model on an RTX 4090 (24 GB)?
Not on a single card at usable quality. Llama 3.3 70B at 4-bit (Q4_K_M) is about 39 GB of weights and ~42–46 GB with context, so it needs roughly 48 GB — typically two 24 GB cards (dual 3090/4090) or an Apple machine with 64 GB+ unified memory. On one 4090, run a 32B model (Qwen2.5 32B or DeepSeek-R1 32B) at 4-bit instead — that fits comfortably.
What's the best LLM for a 16 GB GPU (RTX 4080 / 5080 / 4060 Ti 16GB)?
At 4-bit you comfortably fit 7B–14B models — Llama 3.1 8B, Qwen2.5 14B, Phi-4 14B, Gemma 3 12B — with room for long context. A 24B (Mistral Small 3) is tight but workable at short context; 27B (Gemma) is borderline. Note the 4060 Ti 16GB has a narrow memory bus, so it fits the same models but generates noticeably slower than a 4080/5080.
Does Apple unified memory count as VRAM for LLMs?
Yes, and it's one of Apple Silicon's biggest advantages — the GPU shares the whole unified memory pool. But macOS reserves part of it for the system, so the practical budget for model weights is roughly 70–75% of the total. A 64 GB Mac gives ~48 GB usable (enough for a 70B at 4-bit), and 128 GB gives ~96 GB usable. You can raise the GPU memory cap with a sysctl tweak if you need more.
FP16 vs 8-bit vs 4-bit — which quantization should I use?
For local chat, 4-bit (Q4_K_M) is what almost everyone runs: it cuts VRAM ~75% versus FP16 while losing only ~1.5–2 points on benchmarks. 8-bit is near-lossless (under ~1 point) but doubles the memory of 4-bit. FP16 is reference quality but rarely worth it locally — it's mostly for training/fine-tuning, not inference.
Why does longer context need more VRAM?
Beyond the fixed weight memory, the model stores a KV cache for every token in the conversation, and that grows with context length. A short 4K chat adds only ~12% overhead, but a full 128K-token context can add 50–60% on top of the weights — which can push a model that 'fits' at 8K into 'won't fit' at 128K.