Tag
#vram
Every story tagged vram, newest first.

Which local LLMs fit in 8, 12, 16, 24 or 32GB of VRAM
At Q4_K_M and 8K context, 8GB comfortably holds Gemma 2 9B (5.8 GB) and 24GB holds Gemma 3 27B (17.5 GB). Full tables for every common VRAM tier, computed the same way our GPU checker computes them.
Ahmad J · Sep 6, 2026 · 8 min read

How Much VRAM You Actually Need to Run a Local LLM
VRAM is the hard constraint on running a local LLM. Here's the real math: parameters, precision, quantization, KV cache, what fits on 8GB, 24GB, 32GB, and unified-memory machines, plus where quality and speed actually break.
Ahmad J · Aug 5, 2026 · 7 min read

The best local LLM runners in 2026: Ollama, LM Studio, vLLM, and more
For most people the best local LLM runner in 2026 is still Ollama: free, cross-platform, out of your way. But LM Studio, vLLM, Apple MLX, llama.cpp, Jan, GPT4All, and Open WebUI each win a specific job. Here's which to pick, and what actually fits your GPU.
Ahmad J · Aug 4, 2026 · 9 min read

The best home-server hardware for self-hosting AI in 2026
The used RTX 3090 is still the value pick for local AI in 2026, but a brutal memory shortage reshuffled every price, and 128GB unified boxes (DGX Spark, Strix Halo, Mac Studio) now run big MoE models a 24GB GPU can't hold. Here's what to actually buy.
Ahmad J · Aug 3, 2026 · 13 min read

The best budget laptops for programming and AI work in 2026
The best budget laptop for programming and AI in 2026 is the MacBook Air M4. For CUDA work, the Acer Nitro V16 AI: RTX 5050, 32GB RAM, 1TB.
Ahmad J · Jun 20, 2026 · 13 min read
